4 papers
DREAMSTEER: Latent World Models Can Steer VLA Policies During Deployment Without Any Finetuning
Hanchen Cui, Sergio Arnaud, Arjun Majumdar +5
Pretrained vision-language-action (VLA) policies show promising zero-shot generalization, but often fail under deployment-time distribution shift, leading to decreased robustness a…
WorldPlanner: Monte Carlo Tree Search and MPC with Action-Conditioned Visual World Models
R. Khorrambakht, Joaquim Ortiz-Haro, Joseph Amigo +4
Robots must understand their environment from raw sensory inputs and reason about the consequences of their actions in it to solve complex tasks. Behavior Cloning (BC) leverages ta…
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Mido Assran, Adrien Bardes, David Fan +27
A major challenge for modern AI is to learn to understand the world and learn to act largely by observation. This paper explores a self-supervised approach that combines internet-s…
Locate 3D: Real-World Object Localization via Self-Supervised Learning in 3D
Sergio Arnaud, Paul McVay, Ada Martin +19
We present LOCATE 3D, a model for localizing objects in 3D scenes from referring expressions like "the small coffee table between the sofa and the lamp." LOCATE 3D sets a new state…