3 papers
cs.CV2026
Dual-Stream Cross-Anchor Correction Grounding Long-Form Captions and the Domain Limits of Object-Level Anchors
Lingkai Bu, Qian Gao, Jun Fan +4
Object hallucination in multimodal large language models arises when language priors and corpus co-occurrence bias outweigh the visual evidence, with nothing tying an object mentio…
cs.CV2026
Diagnosing Dense Same-Class Attribute Misbinding in Large Vision-Language Models
Yuanzhi Xu, Qian Gao, Jun Fan +4
Large vision-language models can recognize the objects and attributes in a crowded scene yet assign an attribute to the wrong same-class instance. Generic visual-question-answering…
cs.AI2026
Mitigating Object Hallucinations in Vision-Language Models through Region-Aware Attention Recalibration
Yuanzhi Xu, Qian Gao, Jun Fan +4
The generation of factually incorrect objects, commonly known as object hallucination, remains a persistent challenge in Large Vision-Language Models (LVLMs). Current approaches to…