activity
20242026
collaborators

21 papers

cs.CV2026

OvisOCR2 Technical Report

Shiyin Lu, Yinglun Li, Yu Xia +10

OvisOCR2 is a 0.8 B parameter end‑to‑end model that converts document page images into Markdown, handling text, formulas, tables, and visual regions, and achieves state‑of‑the‑art…

cs.CL2026

Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding

Sensen Gao, Shanshan Zhao, Xu Jiang +7

Document understanding is critical for applications from financial analysis to scientific discovery. Current approaches, whether OCR-based pipelines feeding Large Language Models (…

cs.CV2026

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization

Wenhao Yang, Yu Xia, Jinlong Huang +10

Recent advancements in Multimodal Large Language Models (MLLMs) have incentivized models to ``think with images'' by actively invoking visual tools during multi-turn reasoning. The…

cs.AI2026

Building Autonomous GUI Navigation via Agentic-Q Estimation and Step-Wise Policy Optimization

Yibo Wang, Guangda Huzhang, Yuwei Hu +7

Recent advances in Multimodal Large Language Models (MLLMs) have substantially driven the progress of autonomous agents for Graphical User Interface (GUI). Nevertheless, in real-wo…

cs.CV2026

Omni-View: Unlocking How Generation Facilitates Understanding in Unified 3D Model based on Multiview images

JiaKui Hu, Shanshan Zhao, Qing-Guo Chen +6

This paper presents Omni-View, which extends the unified multimodal understanding and generation to 3D scenes based on multiview images, exploring the principle that "generation fa…

cs.CV2026

Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities

Shanshan Zhao, Xinjie Zhang, Jintao Guo +9

Recent years have seen remarkable progress in both multimodal understanding models and image generation models. Despite their respective successes, these two domains have evolved i…