8 papers
SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models
Zonghe Liu, Shanyuan Jie, Xiaoquan Sun +4
Vision-Language-Action (VLA) models have shown strong potential for general robot manipulation, but most existing models rely on 2D visual-language backbones and lack fine-grained…
HiMem-WAM: Hierarchical Memory-Gated World Action Models for Robotic Manipulation
Xiaoquan Sun, Ruijian Zhang, Chen Cao +12
World Action Models (WAMs) have emerged as a new powerful paradigm for embodied intelligence, learning action-relevant visual dynamics that significantly enhance generalization and…
LargeMonitor: Monitoring Online Task-Free Continual Learning via Large Pretrained Models
Mingqi Yuan, Xiaoquan Sun, Shihao Luo +1
Online task-free continual learning (TFCL) requires intelligent agents to sequentially accumulate knowledge from an unbounded, non-stationary data stream under strict single-pass c…
Can Vision-Language-Action Models Learn from Real-World Data Continually without Forgetting?
Jiarun Zhu, Yijun Hong, Xiaoquan Sun +7
Vision-Language-Action (VLA) models provide a promising foundation for general-purpose robotics, yet their real-world deployment demands the ability to continually acquire new skil…
AtomVLA: Scalable Post-Training for Robotic Manipulation via Predictive Latent World Models
Xiaoquan Sun, Zetian Xu, Chen Cao +9
Vision-Language-Action (VLA) models demonstrate remarkable potential for generalizable robotic manipulation. The execution of complex multi-step behaviors in VLA models can be impr…
Rethinking the Practicality of Vision-language-action Model: A Comprehensive Benchmark and An Improved Baseline
Wenxuan Song, Jiayi Chen, Xiaoquan Sun +12
Vision-Language-Action (VLA) models have emerged as a generalist robotic agent. However, existing VLAs are hindered by excessive parameter scales, prohibitive pre-training requirem…