collaborators

6 papers

cs.RO2026

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling

Yihan Lin, Jiawei He, Shifeng Bao +6

Robust robot control benefits from explicitly modeling state transitions, but video-generation world action models (WAMs) introduce substantial deployment cost. Existing latent WAM…

cs.RO2026

Training Vision-Language-Action Models with Dense Embodied Chain-of-Thought Supervision

Haoyang Li, Guanlin Li, Youhe Feng +9

Cross-embodiment transfer in vision-language-action (VLA) models remains challenging because low-level state and action spaces differ fundamentally across robot platforms. We obser…

cs.RO2026

From Pixels to Tokens: A Systematic Study of Latent Action Supervision for Vision-Language-Action Models

Yihan Lin, Haoyang Li, Yang Li +4

Latent actions serve as an intermediate representation that enables consistent modeling of vision-language-action (VLA) models across heterogeneous datasets. However, approaches to…

cs.CV2026

Action Draft and Verify: A Self-Verifying Framework for Vision-Language-Action Model

Chen Zhao, Zhuoran Wang, Haoyang Li +6

Vision-Language-Action (VLA) models have recently demonstrated strong performance across embodied tasks. Modern VLAs commonly employ diffusion action experts to efficiently generat…

cs.CV2025

Ovis2.5 Technical Report

Shiyin Lu, Yang Li, Yu Xia +39

We present Ovis2.5, a successor to Ovis2 designed for native-resolution visual perception and strong multimodal reasoning. Ovis2.5 integrates a native-resolution vision transformer…

cs.CV2025

Ovis-U1 Technical Report

Guo-Hua Wang, Shanshan Zhao, Xinjie Zhang +9

In this report, we introduce Ovis-U1, a 3-billion-parameter unified model that integrates multimodal understanding, text-to-image generation, and image editing capabilities. Buildi…