4 papers
Global Context or Local Detail? Adaptive Visual Grounding for Hallucination Mitigation
Yubo Jiang, Xin Yang, Abudukelimu Wuerkaixi +7
Vision-Language Models (VLMs) are frequently undermined by object hallucination--generating content that contradicts visual reality--due to an over-reliance on linguistic priors. W…
Breaking the Illusion: When Positive Meets Negative in Multimodal Decoding
Yubo Jiang, Yitong An, Xin Yang +7
Vision-Language Models (VLMs) are frequently undermined by object hallucination, generating content that contradicts visual reality, due to an over-reliance on linguistic priors. W…
AutothinkRAG: Complexity-Aware Control of Retrieval-Augmented Reasoning for Image-Text Interaction
Jiashu Yang, Chi Zhang, Abudukelimu Wuerkaixi +5
Multimodal document question answering requires retrieving dispersed evidence from visually rich long documents and performing reliable reasoning over heterogeneous information. Ex…
MMRA: A Benchmark for Evaluating Multi-Granularity and Multi-Image Relational Association Capabilities in Large Visual Language Models
Siwei Wu, Kang Zhu, Yu Bai +10
Given the remarkable success that large visual language models (LVLMs) have achieved in image perception tasks, the endeavor to make LVLMs perceive the world like humans is drawing…