causal pretraining 1foundation models 1robot control 1sparse mixture of experts 1video-action models 1
From the 1 of 16 linked papers with an AI index.
Showing cs.ROShow all
2 papers · 1 filter
cs.RO2026
Native Video-Action Pretraining for Generalizable Robot Control
Qihang Zhang, Lin Li, Luyao Zhang +26
The paper introduces LingBot-VA 2.0, a video-action foundation model designed specifically for robot control, featuring a semantic visual-action tokenizer, causal pretraining, a sp…
cs.RO2026
ThinkingVLA: Interleaved Vision and Language Reasoning for Robotic Manipulation
Tianyi Lu, Hui Zhang, Zijie Diao +8
Most Vision-Language-Action (VLA) models map observations directly to actions without explicit reasoning, limiting their capacity for reasoning-intensive long-horizon tasks. To add…