activity
20242026
collaborators

8 papers

cs.CL2026

D2-ScaleAgent: Dual-Dimensional Scaling for Long Document Understanding

Hao Zhang, Longrong Yang, Lunhao Duan +3

Multi-modal retrieval-augmented generation (RAG) is a key technique for visually rich long document understanding. Existing multi-modal RAG methods are progressively advancing towa…

cs.CL2026

Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding

Sensen Gao, Shanshan Zhao, Xu Jiang +7

Document understanding is critical for applications from financial analysis to scientific discovery. Current approaches, whether OCR-based pipelines feeding Large Language Models (…

cs.CV2026

Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities

Shanshan Zhao, Xinjie Zhang, Jintao Guo +9

Recent years have seen remarkable progress in both multimodal understanding models and image generation models. Despite their respective successes, these two domains have evolved i…

cs.CV2025

Ovis2.5 Technical Report

Shiyin Lu, Yang Li, Yu Xia +39

We present Ovis2.5, a successor to Ovis2 designed for native-resolution visual perception and strong multimodal reasoning. Ovis2.5 integrates a native-resolution vision transformer…

cs.CV2025

Harnessing Text-to-Image Diffusion Models for Point Cloud Self-Supervised Learning

Yiyang Chen, Shanshan Zhao, Lunhao Duan +2

Diffusion-based models, widely used in text-to-image generation, have proven effective in 2D representation learning. Recently, this framework has been extended to 3D self-supervis…

cs.CV2025

Ovis-U1 Technical Report

Guo-Hua Wang, Shanshan Zhao, Xinjie Zhang +9

In this report, we introduce Ovis-U1, a 3-billion-parameter unified model that integrates multimodal understanding, text-to-image generation, and image editing capabilities. Buildi…