5 papers
FARM: Find Anything using Relational Spatial Memory
Siming He, Leo Huang, Adam Lilja +7
Robots operating in homes, warehouses, and other object-rich environments need memory systems that can find specific object instances on demand. Object-level memory alone is often…
T-Rex: Tactile-Reactive Dexterous Manipulation
Dantong Niu, Zhuoyang Liu, Zekai Wang +31
The ability to react dynamically to tactile signals has long been considered crucial to agile human-level dexterity. Yet contemporary learning-based Vision-Language-Action (VLA) mo…
Forecasting Motion in the Wild
Neerja Thakkar, Shiry Ginosar, Jacob Walker +3
Visual intelligence requires anticipating the future behavior of agents, yet vision systems lack a general representation for motion and behavior. We propose dense point trajectori…
Whole-Body Conditioned Egocentric Video Prediction
Yutong Bai, Danny Tran, Amir Bar +3
We train models to Predict Ego-centric Video from human Actions (PEVA), given the past video and an action represented by the relative 3D body pose. By conditioning on kinematic po…
TARDIS STRIDE: A Spatio-Temporal Road Image Dataset and World Model for Autonomy
Héctor Carrión, Yutong Bai, VÃctor A. Hernández Castro +6
World models aim to simulate environments and enable effective agent behavior. However, modeling real-world environments presents unique challenges as they dynamically change acros…