collaborators

16 papers

cs.AI2026

PaDoc: Layout-Grounded Parallel Decoding for Document Parsing

Hao Yu, Jiabo Zhan, Kang Liu +8

End-to-end document parsers provide a unified interface, but serialize page layouts and regional contents into one autoregressive sequence. This formulation forces independent regi…

cs.CV2026

Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings

Peixi Wu, Ke Mei, Feipeng Ma +15

The paper introduces RIME, a rewrite-driven framework that improves multimodal embeddings by jointly optimizing generation and retrieval-friendly rewriting, aligning generative and…

cs.CV2026

MMAgent-R: Learning to Rerank and Reject for Agentic mRAG

Tao Zhang, Ziqi Zhang, Zongyang Ma +7

Knowledge-based Visual Question Answering (KB-VQA) requires models to retrieve visual entities matching the query image from large-scale encyclopedic knowledge bases and answer rel…

cs.CV2026

Segmentation before Answering: Pixel Grounding for MLLM Visual Reasoning

Yake Wei, Yuan Wang, Fengyun Rao +2

Recent advancements in Multimodal Large Language Models (MLLMs) have evolved from static perception to interleaved visual-language reasoning, often referred to as ``thinking with i…

cs.CV2026

ObjEmbed: Towards Universal Multimodal Object Embeddings

Shenghao Fu, Yukun Su, Fengyun Rao +3

Aligning objects with corresponding textual descriptions is a fundamental challenge and a realistic requirement in vision-language understanding. While recent multimodal embedding…

cs.CV2026

Semantic-Enriched Latent Visual Reasoning

Tianrun Xu, Yue Sun, Qixun Wang +8

Multimodal latent-space reasoning aims to replace explicit thinking with images by performing visual reasoning directly in a compact latent space. However, existing approaches larg…