From the 1 of 12 linked papers with an AI index.
12 papers
Native Video-Action Pretraining for Generalizable Robot Control
Qihang Zhang, Lin Li, Luyao Zhang +26
The paper introduces LingBot-VA 2.0, a video-action foundation model designed specifically for robot control, featuring a semantic visual-action tokenizer, causal pretraining, a sp…
RepWAM: World Action Modeling with Representation Visual-Action Tokenizers
Junke Wang, Qihang Zhang, Shuai Yang +5
This work presents RepWAM, a representation-centric world action model (WAM) built on representation visual-action tokenizers. Existing WAMs typically inherit reconstruction-orient…
A Progressive Training Strategy for Vision-Language Models to Counteract Spatio-Temporal Hallucinations in Embodied Reasoning
Xiaoda Yang, Shuai Yang, Can Wang +9
Vision-Language Models (VLMs) have made significant strides in static image understanding but continue to face critical hurdles in spatiotemporal reasoning. A major bottleneck is "…
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
Shuai Yang, Hao Li, Bin Wang +7
To operate effectively in the real world, robots should integrate multimodal reasoning with precise action generation. However, existing vision-language-action (VLA) models often s…
RoboInter: A Holistic Intermediate Representation Suite Towards Robotic Manipulation
Hao Li, Ziqin Wang, Zi-han Ding +9
Advances in large vision-language models (VLMs) have stimulated growing interest in vision-language-action (VLA) systems for robot manipulation. However, existing manipulation data…
CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling
Hao Li, Shuai Yang, Yilun Chen +8
Recent vision-language-action (VLA) models built on pretrained vision-language models (VLMs) have demonstrated strong performance in robotic manipulation. However, these models rem…