16 papers
PaDoc: Layout-Grounded Parallel Decoding for Document Parsing
Hao Yu, Jiabo Zhan, Kang Liu +8
End-to-end document parsers provide a unified interface, but serialize page layouts and regional contents into one autoregressive sequence. This formulation forces independent regi…
Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings
Peixi Wu, Ke Mei, Feipeng Ma +15
The paper introduces RIME, a rewrite-driven framework that improves multimodal embeddings by jointly optimizing generation and retrieval-friendly rewriting, aligning generative and…
MMAgent-R: Learning to Rerank and Reject for Agentic mRAG
Tao Zhang, Ziqi Zhang, Zongyang Ma +7
Knowledge-based Visual Question Answering (KB-VQA) requires models to retrieve visual entities matching the query image from large-scale encyclopedic knowledge bases and answer rel…
Segmentation before Answering: Pixel Grounding for MLLM Visual Reasoning
Yake Wei, Yuan Wang, Fengyun Rao +2
Recent advancements in Multimodal Large Language Models (MLLMs) have evolved from static perception to interleaved visual-language reasoning, often referred to as ``thinking with i…
ObjEmbed: Towards Universal Multimodal Object Embeddings
Shenghao Fu, Yukun Su, Fengyun Rao +3
Aligning objects with corresponding textual descriptions is a fundamental challenge and a realistic requirement in vision-language understanding. While recent multimodal embedding…
Semantic-Enriched Latent Visual Reasoning
Tianrun Xu, Yue Sun, Qixun Wang +8
Multimodal latent-space reasoning aims to replace explicit thinking with images by performing visual reasoning directly in a compact latent space. However, existing approaches larg…