3 papers
cs.CV2026
V-Zero: Self-Improving Multimodal Reasoning with Zero Annotation
Han Wang, Yi Yang, Jingyuan Hu +2
Recent advances in multimodal learning have significantly enhanced the reasoning capabilities of vision-language models (VLMs). However, state-of-the-art approaches rely heavily on…
cs.CV2025
R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization
Yi Yang, Xiaoxuan He, Hongkun Pan +9
Large Language Models have demonstrated remarkable reasoning capability in complex textual tasks. However, multimodal reasoning, which requires integrating visual and textual infor…
cs.AI2024
Exploring the Interplay Between Video Generation and World Models in Autonomous Driving: A Survey
Ao Fu, Yi Zhou, Tao Zhou +5
World models and video generation are pivotal technologies in the domain of autonomous driving, each playing a critical role in enhancing the robustness and reliability of autonomo…