14 papers
MMAgent-R: Learning to Rerank and Reject for Agentic mRAG
Tao Zhang, Ziqi Zhang, Zongyang Ma +7
Knowledge-based Visual Question Answering (KB-VQA) requires models to retrieve visual entities matching the query image from large-scale encyclopedic knowledge bases and answer rel…
How Far Is Document Parsing from Solved? PureDocBench: A Source-TraceableBenchmark across Clean, Degraded, and Real-World Settings
Zhiheng Li, Zongyang Ma, Jiaxian Chen +12
The past year has seen over 20 open-source document parsing models, yet thefield still benchmarks almost exclusively on OmniDocBench, a 1,355-pagemanually annotated dataset whose t…
Token Caching for Diffusion Transformer Acceleration
Jinming Lou, Wenyang Luo, Yufan Liu +5
Diffusion transformers have gained substantial interest in diffusion generative modeling due to their outstanding performance. However, their computational demands, particularly th…
MMhops-R1: Multimodal Multi-hop Reasoning
Tao Zhang, Ziqi Zhang, Zongyang Ma +7
The ability to perform multi-modal multi-hop reasoning by iteratively integrating information across various modalities and external knowledge is critical for addressing complex re…
Reversing Flow for Image Restoration
Haina Qin, Wenyang Luo, Libin Wang +5
Image restoration aims to recover high-quality (HQ) images from degraded low-quality (LQ) ones by reversing the effects of degradation. Existing generative models for image restora…
Visual-Instructed Degradation Diffusion for All-in-One Image Restoration
Wenyang Luo, Haina Qin, Zewen Chen +6
Image restoration tasks like deblurring, denoising, and dehazing usually need distinct models for each degradation type, restricting their generalization in real-world scenarios wi…