3 papers
cs.CV2026
Unified Video-Action Joint Denoising for Dexterous Action and Data Generation
Dingrui Wang, YuAn Wang, Jinkun Liu +4
Recent world action models leverage video foundation models by aligning broad visual-dynamics priors with executable robot actions. We revisit this alignment from a distributional…
cs.RO2026
OneVLA: A Unified Framework for Embodied Tasks
Lingfeng Zhang, Xiaoshuai Hao, Yingbo Tang +10
Navigation and manipulation are fundamental capabilities of embodied intelligence, enabling robots to interpret natural language commands and interact physically with their surroun…
cs.AI2026
Thinking in Text and Images: Interleaved Vision--Language Reasoning Traces for Long-Horizon Robot Manipulation
Jinkun Liu, Haohan Chi, Lingfeng Zhang +6
Long-horizon robotic manipulation requires plans that are both logically coherent and geometrically grounded. Existing Vision-Language-Action policies usually hide planning in late…