6 papers
JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling
Yihan Lin, Jiawei He, Shifeng Bao +6
Robust robot control benefits from explicitly modeling state transitions, but video-generation world action models (WAMs) introduce substantial deployment cost. Existing latent WAM…
Training Vision-Language-Action Models with Dense Embodied Chain-of-Thought Supervision
Haoyang Li, Guanlin Li, Youhe Feng +9
Cross-embodiment transfer in vision-language-action (VLA) models remains challenging because low-level state and action spaces differ fundamentally across robot platforms. We obser…
From Pixels to Tokens: A Systematic Study of Latent Action Supervision for Vision-Language-Action Models
Yihan Lin, Haoyang Li, Yang Li +4
Latent actions serve as an intermediate representation that enables consistent modeling of vision-language-action (VLA) models across heterogeneous datasets. However, approaches to…
Action Draft and Verify: A Self-Verifying Framework for Vision-Language-Action Model
Chen Zhao, Zhuoran Wang, Haoyang Li +6
Vision-Language-Action (VLA) models have recently demonstrated strong performance across embodied tasks. Modern VLAs commonly employ diffusion action experts to efficiently generat…
Ovis2.5 Technical Report
Shiyin Lu, Yang Li, Yu Xia +39
We present Ovis2.5, a successor to Ovis2 designed for native-resolution visual perception and strong multimodal reasoning. Ovis2.5 integrates a native-resolution vision transformer…
Ovis-U1 Technical Report
Guo-Hua Wang, Shanshan Zhao, Xinjie Zhang +9
In this report, we introduce Ovis-U1, a 3-billion-parameter unified model that integrates multimodal understanding, text-to-image generation, and image editing capabilities. Buildi…