1 citations · 1 across the 2 of their papers we have counts for
15 papers
MJEPA: A Simple and Scalable Joint-Embedding Predictive Architecture for Audio-Visual Learning
Revant Teotia, Adrien Bardes, Michael Rabbat +3
Self-supervised learning from large-scale video data has emerged as a dominant paradigm for visual representation learning. Since audio and visual streams naturally co-occur in vid…
Hierarchical Planning with Latent World Models
Wancong Zhang, Basile Terver, Artem Zholus +8
World models are a promising path to zero-shot embodied control through planning. However, existing world model planners struggle on long-horizon, multi-stage tasks: prediction err…
V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
Lorenzo Mur-Labadia, Matthew Muckley, Amir Bar +6
We present V-JEPA 2.1, a family of self-supervised models that learn dense, high-quality visual representations for both images and videos while retaining strong global scene under…
Beyond Language Modeling: An Exploration of Multimodal Pretraining
Shengbang Tong, David Fan, John Nguyen +18
The visual world offers a critical axis for advancing foundation models beyond language. Despite growing interest in this direction, the design space for native multimodal models r…
Inference-time Physics Alignment of Video Generative Models with Latent World Models
Jianhao Yuan, Xiaofeng Zhang, Felix Friedrich +7
State-of-the-art video generative models produce promising visual content yet often violate basic physics principles, limiting their utility. While some attribute this deficiency t…
Learning Latent Action World Models In The Wild
Quentin Garrido, Tushar Nagarajan, Basile Terver +3
Agents capable of reasoning and planning in the real world require the ability of predicting the consequences of their actions. While world models possess this capability, they mos…