4 papers
GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation
AgiBot Research Team, Renhang Liu, Wenzhi Zhao +42
World-action models (WAM) predict future states to guide robot actions, enabling learning from both action-free video and action-labeled interaction. Most inherit pretrained video…
OmniReasoner: Thinking with Long Audio-Video via Native Tool Use
Yu Chen, Caorui Li, Ziyu Xiong +8
Long audio-video reasoning is difficult for omnimodal LLMs because the decisive evidence is often sparse, cross-modal, and too expensive to preserve with uniformly high-fidelity in…
SOP: A Scalable Online Post-Training System for Vision-Language-Action Models
Mingjie Pan, Siyuan Feng, Qinglin Zhang +9
Vision-language-action (VLA) models achieve strong generalization through large-scale pre-training, but real-world deployment requires expert-level task proficiency in addition to…
ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models
Linqing Zhong, Yi Liu, Yifei Wei +4
Vision-Language-Action models have emerged as essential generalist robot policies for diverse manipulation tasks, conventionally relying on directly translating multimodal inputs i…