collaborators

6 papers

cs.RO2026

DualWAM: Dual-System World Action Models for Asynchronous Global Planning and Local Refinement

Yixin Zheng, Jiangran Lyu, Yuntian Deng +6

World Action Models (WAMs) jointly generate robot actions and predict future world states, transferring priors from video pretraining to robot control. However, future visual predi…

cs.CV2026

World-Ego Modeling for Long-Horizon Evolution in Hybrid Embodied Tasks

Zuyao Lin, Jianhui Zhang, Peidong Jia +3

World models are widely explored in embodied intelligence, yet they typically predict distinct evolutions of the world and the ego within a single stream, where the world captures…

cs.RO2026

Emerging Extrinsic Dexterity in Cluttered Scenes via Dynamics-aware Policy Learning

Yixin Zheng, Jiangran Lyu, Yifan Zhang +8

Extrinsic dexterity leverages environmental contact to overcome the limitations of prehensile manipulation. However, achieving such dexterity in cluttered scenes remains challengin…

cs.RO2026

Reshaping Action Error Distributions for Reliable Vision-Language-Action Models

Shuanghao Bai, Dakai Wang, Cheng Chi +8

In robotic manipulation, vision-language-action (VLA) models have emerged as a promising paradigm for learning generalizable and scalable robot policies. Most existing VLA framewor…

cs.RO2026

Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models

Shuanghao Bai, Jing Lyu, Wanqi Zhou +9

Vision-Language-Action (VLA) models benefit from chain-of-thought (CoT) reasoning, but existing approaches incur high inference overhead and rely on discrete reasoning representati…

cs.RO2025

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision

Shengliang Deng, Mi Yan, Yixin Zheng +7

While Vision-Language-Action (VLA) models excel in generalist manipulation, they often lack fine-grained spatial awareness and show limited viewpoint robustness. This limitation la…