Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Action Draft and Verify: A Self-Verifying Framework for Vision-Language-Action Model
Chen Zhao, Zhuoran Wang, Haoyang Li +6
Vision-Language-Action (VLA) models have recently demonstrated strong performance across embodied tasks. Modern VLAs commonly employ diffusion action experts to efficiently generat…
cs.CV2025
Ovis2.5 Technical Report
Shiyin Lu, Yang Li, Yu Xia +39
We present Ovis2.5, a successor to Ovis2 designed for native-resolution visual perception and strong multimodal reasoning. Ovis2.5 integrates a native-resolution vision transformer…
cs.CV2025
Ovis-U1 Technical Report
Guo-Hua Wang, Shanshan Zhao, Xinjie Zhang +9
In this report, we introduce Ovis-U1, a 3-billion-parameter unified model that integrates multimodal understanding, text-to-image generation, and image editing capabilities. Buildi…