9 papers
IOI: Decoupling Kinematics and Physics for Interactive World Models
Chengyu Bai, Peidong Jia, Tiecheng Guo +11
Developing generalist embodied agents requires interactive environments providing visually realistic feedback and accurate action-conditioned dynamics. Interactive world models add…
MV-WAM: Manifold-Aware World Action Model with Value Augmentation
Jintao Chen, Peidong Jia, Qingpo Wuwu +13
Achieving robust and generalizable manipulation across diverse environments remains a fundamental challenge in embodied robotics. Recent world action models achieve strong in-domai…
EvolveNav: Proactive Preflection and Self-Evolving Memory for Zero-Shot Object Goal Navigation
Qi Chai, Wenhao Shen, Nanjie Yao +5
Zero-Shot Object-Goal Navigation (ZS-OGN) requires embodied agents to explore and locate target objects without any prior training. To this end, recent methods leverage foundation…
Offline Policy Evaluation for Manipulation Policies via Discounted Liveness Formulation
Hao Wang, Joshua Bowden, Colton Crosby +1
Policy evaluation is a fundamental component of the development and deployment pipeline for robotic policies. In modern manipulation systems, this problem is particularly challengi…
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models
Hao Wang, Xiaobao Wei, Jingyang He +10
Precise spatial reasoning is fundamental to robotic manipulation, yet the visual backbones of current vision-language-action (VLA) models are predominantly pretrained on 2D image d…
LoopVLA: Learning Sufficiency in Recurrent Refinement for Vision-Language-Action Models
Boyang Shen, Kaixiang Yang, Hao Wang +4
Current Vision-Language-Action (VLA) models typically treat the deepest representation of a vision-language backbone as universally optimal for action prediction. However, robotic…