2 papers
cs.RO2026
SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models
Zonghe Liu, Shanyuan Jie, Xiaoquan Sun +4
Vision-Language-Action (VLA) models have shown strong potential for general robot manipulation, but most existing models rely on 2D visual-language backbones and lack fine-grained…
cs.RO2026
AtomVLA: Scalable Post-Training for Robotic Manipulation via Predictive Latent World Models
Xiaoquan Sun, Zetian Xu, Chen Cao +9
Vision-Language-Action (VLA) models demonstrate remarkable potential for generalizable robotic manipulation. The execution of complex multi-step behaviors in VLA models can be impr…