collaborators

6 papers

cs.CV2026

Plan, Watch, Recover: A Benchmark and Architectures for Proactive Procedural Assistance

Kaustav Kundu, Ritvik Shrivastava, Maxim Arap +13

We envision a proactive multi-modal assistant system which gives users real-time step-by-step guidance on a procedural task, autonomously deciding \textit{when} to interrupt, and \…

cs.SE2026

Computer Use at the Edge of the Statistical Precipice

Pierluca D'Oro, Sneha Silwal, William Wong +6

Evaluating Computer Use Agents (CUAs) on interactive environments is fraught with methodological pitfalls that the field has yet to systematically address. We show that a 1MB repla…

cs.CV2026

VL-JEPA: Joint Embedding Predictive Architecture for Vision-language

Delong Chen, Mustafa Shukor, Theo Moutakanni +7

We introduce VL-JEPA, a vision-language model built on a Joint Embedding Predictive Architecture (JEPA). Instead of autoregressively generating tokens as in classical VLMs, VL-JEPA…

cs.CV2026

Action100M: A Large-scale Video Action Dataset

Delong Chen, Tejaswi Kasarla, Yejin Bang +6

Inferring physical actions from visual observations is a fundamental capability for advancing machine intelligence in the physical world. Achieving this requires large-scale, open-…

cs.AI2025

DigiData: Training and Evaluating General-Purpose Mobile Control Agents

Yuxuan Sun, Manchen Wang, Shengyi Qian +18

AI agents capable of controlling user interfaces have the potential to transform human interaction with digital devices. To accelerate this transformation, two fundamental building…

cs.AI2025

Planning with Reasoning using Vision Language World Model

Delong Chen, Theo Moutakanni, Willy Chung +4

Effective planning requires strong world models, but high-level world models that can understand and reason about actions with semantic and temporal abstraction remain largely unde…