6 papers
DREAMSTEER: Latent World Models Can Steer VLA Policies During Deployment Without Any Finetuning
Hanchen Cui, Sergio Arnaud, Arjun Majumdar +5
Pretrained vision-language-action (VLA) policies show promising zero-shot generalization, but often fail under deployment-time distribution shift, leading to decreased robustness a…
Visuo-Tactile World Models
Carolina Higuera, Sergio Arnaud, Byron Boots +3
We introduce multi-task Visuo-Tactile World Models (VT-WM), which capture the physics of contact through touch reasoning. By complementing vision with tactile sensing, VT-WM better…
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Mido Assran, Adrien Bardes, David Fan +27
A major challenge for modern AI is to learn to understand the world and learn to act largely by observation. This paper explores a self-supervised approach that combines internet-s…
From Thousands to Billions: 3D Visual Language Grounding via Render-Supervised Distillation from 2D VLMs
Ang Cao, Sergio Arnaud, Oleksandr Maksymets +12
3D vision-language grounding faces a fundamental data bottleneck: while 2D models train on billions of images, 3D models have access to only thousands of labeled scenes--a six-orde…
Unifying 2D and 3D Vision-Language Understanding
Ayush Jain, Alexander Swerdlow, Yuzhou Wang +5
Progress in 3D vision-language learning has been hindered by the scarcity of large-scale 3D datasets. We introduce UniVLG, a unified architecture for 2D and 3D vision-language unde…
Locate 3D: Real-World Object Localization via Self-Supervised Learning in 3D
Sergio Arnaud, Paul McVay, Ada Martin +19
We present LOCATE 3D, a model for localizing objects in 3D scenes from referring expressions like "the small coffee table between the sofa and the lamp." LOCATE 3D sets a new state…