5 papers
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
Shengliang Deng, Mi Yan, Yixin Zheng +7
While Vision-Language-Action (VLA) models excel in generalist manipulation, they often lack fine-grained spatial awareness and show limited viewpoint robustness. This limitation la…
Emerging Extrinsic Dexterity in Cluttered Scenes via Dynamics-aware Policy Learning
Yixin Zheng, Jiangran Lyu, Yifan Zhang +8
Extrinsic dexterity leverages environmental contact to overcome the limitations of prehensile manipulation. However, achieving such dexterity in cluttered scenes remains challengin…
World-Ego Modeling for Long-Horizon Evolution in Hybrid Embodied Tasks
Zuyao Lin, Jianhui Zhang, Peidong Jia +3
World models are widely explored in embodied intelligence, yet they typically predict distinct evolutions of the world and the ego within a single stream, where the world captures…
Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models
Shuanghao Bai, Jing Lyu, Wanqi Zhou +9
Vision-Language-Action (VLA) models benefit from chain-of-thought (CoT) reasoning, but existing approaches incur high inference overhead and rely on discrete reasoning representati…
Reshaping Action Error Distributions for Reliable Vision-Language-Action Models
Shuanghao Bai, Dakai Wang, Cheng Chi +8
In robotic manipulation, vision-language-action (VLA) models have emerged as a promising paradigm for learning generalizable and scalable robot policies. Most existing VLA framewor…