7 papers
Who Needs Labels? Adapting Vision Foundation Models With the Metadata You Already Have
Elouan Gardès, Seung Eun Yi, Kartik Ahuja +6
We propose a label-free approach to adapt powerful but generic vision foundation models to specialized scientific domains. Standard supervised fine-tuning is often ill-suited to th…
Plan, Watch, Recover: A Benchmark and Architectures for Proactive Procedural Assistance
Kaustav Kundu, Ritvik Shrivastava, Maxim Arap +13
We envision a proactive multi-modal assistant system which gives users real-time step-by-step guidance on a procedural task, autonomously deciding \textit{when} to interrupt, and \…
VL-JEPA: Joint Embedding Predictive Architecture for Vision-language
Delong Chen, Mustafa Shukor, Theo Moutakanni +7
We introduce VL-JEPA, a vision-language model built on a Joint Embedding Predictive Architecture (JEPA). Instead of autoregressively generating tokens as in classical VLMs, VL-JEPA…
Action100M: A Large-scale Video Action Dataset
Delong Chen, Tejaswi Kasarla, Yejin Bang +6
Inferring physical actions from visual observations is a fundamental capability for advancing machine intelligence in the physical world. Achieving this requires large-scale, open-…
Planning with Reasoning using Vision Language World Model
Delong Chen, Theo Moutakanni, Willy Chung +4
Effective planning requires strong world models, but high-level world models that can understand and reason about actions with semantic and temporal abstraction remain largely unde…
Embodied AI Agents: Modeling the World
Pascale Fung, Yoram Bachrach, Asli Celikyilmaz +18
This paper describes our research on AI agents embodied in visual, virtual or physical forms, enabling them to interact with both users and their environments. These agents, which…