activity
20242026
collaborators

9 papers

cs.CL2026

Reward Prediction with Factorized World States

Yijun Shen, Delong Chen, Xianming Hu +4

Agents must infer action outcomes and select actions that maximize a reward signal indicating how close the goal is to being reached. Supervised learning of reward models could int…

cs.CV2026

VL-JEPA: Joint Embedding Predictive Architecture for Vision-language

Delong Chen, Mustafa Shukor, Theo Moutakanni +7

We introduce VL-JEPA, a vision-language model built on a Joint Embedding Predictive Architecture (JEPA). Instead of autoregressively generating tokens as in classical VLMs, VL-JEPA…

cs.CV2026

Action100M: A Large-scale Video Action Dataset

Delong Chen, Tejaswi Kasarla, Yejin Bang +6

Inferring physical actions from visual observations is a fundamental capability for advancing machine intelligence in the physical world. Achieving this requires large-scale, open-…

cs.LG2025

TV2TV: A Unified Framework for Interleaved Language and Video Generation

Xiaochuang Han, Youssef Emad, Melissa Hall +15

Video generation models are rapidly advancing, but can still struggle with complex video outputs that require significant semantic branching or repeated high-level reasoning about…

cs.AI2025

Planning with Reasoning using Vision Language World Model

Delong Chen, Theo Moutakanni, Willy Chung +4

Effective planning requires strong world models, but high-level world models that can understand and reason about actions with semantic and temporal abstraction remain largely unde…

cs.AI2025

Embodied AI Agents: Modeling the World

Pascale Fung, Yoram Bachrach, Asli Celikyilmaz +18

This paper describes our research on AI agents embodied in visual, virtual or physical forms, enabling them to interact with both users and their environments. These agents, which…