3 papers
cs.CV2026
MoCHA: Denoising Caption Supervision for Motion-Text Retrieval
Nikolai Warner, Cameron Ethan Taylor, Irfan Essa +1
Text-motion retrieval systems learn shared embedding spaces from motion-caption pairs via contrastive objectives. However, each caption is not a deterministic label but a sample fr…
cs.CV2025
Leveraging Procedural Knowledge and Task Hierarchies for Efficient Instructional Video Pre-training
Karan Samel, Nitish Sontakke, Irfan Essa
Instructional videos provide a convenient modality to learn new tasks (ex. cooking a recipe, or assembling furniture). A viewer will want to find a corresponding video that reflect…
cs.CV2024
Exploring Efficient Foundational Multi-modal Models for Video Summarization
Karan Samel, Apoorva Beedu, Nitish Sontakke +1
Foundational models are able to generate text outputs given prompt instructions and text, audio, or image inputs. Recently these models have been combined to perform tasks on video…