From the 1 of 10 linked papers with an AI index.
10 papers
Gaussian Mixture Modeling for Event-Aware Visual Allocation in Long Video Understanding
Yifan Lu, Ziqi Zhang, Chunfeng Yuan +3
The paper introduces GMM-EVA, a training-free framework that uses Gaussian Mixture Models to detect event-level structures in long videos and allocate visual tokens by selecting on…
MMAgent-R: Learning to Rerank and Reject for Agentic mRAG
Tao Zhang, Ziqi Zhang, Zongyang Ma +7
Knowledge-based Visual Question Answering (KB-VQA) requires models to retrieve visual entities matching the query image from large-scale encyclopedic knowledge bases and answer rel…
How Far Is Document Parsing from Solved? PureDocBench: A Source-TraceableBenchmark across Clean, Degraded, and Real-World Settings
Zhiheng Li, Zongyang Ma, Jiaxian Chen +12
The past year has seen over 20 open-source document parsing models, yet thefield still benchmarks almost exclusively on OmniDocBench, a 1,355-pagemanually annotated dataset whose t…
Making MLLMs Blind: Adversarial Smuggling Attacks in MLLM Content Moderation
Zhiheng Li, Zongyang Ma, Yuntong Pan +8
Multimodal Large Language Models (MLLMs) are increasingly being deployed as automated content moderators. Within this landscape, we uncover a critical threat: Adversarial Smuggling…
Beyond Semantic Search: Towards Referential Anchoring in Composed Image Retrieval
Yuxin Yang, Yinan Zhou, Yuxin Chen +6
Composed Image Retrieval (CIR) has demonstrated significant potential by enabling flexible multimodal queries that combine a reference image and modification text. However, CIR inh…
MMhops-R1: Multimodal Multi-hop Reasoning
Tao Zhang, Ziqi Zhang, Zongyang Ma +7
The ability to perform multi-modal multi-hop reasoning by iteratively integrating information across various modalities and external knowledge is critical for addressing complex re…