collaborators

5 papers

cs.RO2026

SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models

Zonghe Liu, Shanyuan Jie, Xiaoquan Sun +4

Vision-Language-Action (VLA) models have shown strong potential for general robot manipulation, but most existing models rely on 2D visual-language backbones and lack fine-grained…

cs.RO2026

HiMem-WAM: Hierarchical Memory-Gated World Action Models for Robotic Manipulation

Xiaoquan Sun, Ruijian Zhang, Chen Cao +12

World Action Models (WAMs) have emerged as a new powerful paradigm for embodied intelligence, learning action-relevant visual dynamics that significantly enhance generalization and…

cs.RO2026

Can Vision-Language-Action Models Learn from Real-World Data Continually without Forgetting?

Jiarun Zhu, Yijun Hong, Xiaoquan Sun +7

Vision-Language-Action (VLA) models provide a promising foundation for general-purpose robotics, yet their real-world deployment demands the ability to continually acquire new skil…

cs.RO2026

AIM: Intent-Aware Unified world action Modeling with Spatial Value Maps

Liaoyuan Fan, Zetian Xu, Chen Cao +3

Pretrained video generation models provide strong priors for robot control, but existing unified world action models still struggle to decode reliable actions without substantial r…

cs.RO2026

AtomVLA: Scalable Post-Training for Robotic Manipulation via Predictive Latent World Models

Xiaoquan Sun, Zetian Xu, Chen Cao +9

Vision-Language-Action (VLA) models demonstrate remarkable potential for generalizable robotic manipulation. The execution of complex multi-step behaviors in VLA models can be impr…