4 papers
Simple Recipe Works: Vision-Language-Action Models are Natural Continual Learners with Reinforcement Learning
Jiaheng Hu, Jay Shim, Chen Tang +4
Continual Reinforcement Learning (CRL) for Vision-Language-Action (VLA) models is a promising direction toward self-improving embodied agents that can adapt in openended, evolving…
EgoExo-WM: Unlocking Exo Video for Ego World Models
Danny Tran, Roberto MartÃn-MartÃn, Kristen Grauman
Egocentric world models present a promising direction for enabling agents to predict and plan, but their performance is constrained by the limited availability of egocentric traini…
Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models
Yurou Yang, Muyuan Lin, Roberto Martin-Martin +4
Recent work explores new opportunities at the intersection of vision-language-action models (VLAs) and geometric foundation models (GFMs) for 3D reconstruction, such as VGGT. While…
Mash, Spread, Slice! Learning to Manipulate Object States via Visual Spatial Progress
Priyanka Mandikal, Jiaheng Hu, Shivin Dass +3
Most robot manipulation focuses on changing the kinematic state of objects: picking, placing, opening, or rotating them. However, a wide range of real-world manipulation tasks invo…