2 papers
cs.CV2026
ActiveScope: Actively Seeking and Correcting Perception for MLLMs
Yajing Wang, Chao Bi, Junshu Sun +4
Multimodal Large Language Models (MLLMs) have demonstrated impressive vision-language understanding, yet still struggle with fine-grained perception in high-resolution images. Whil…
cs.CV2026
Locate-then-Sparsify: Attribution Guided Sparse Strategy for Visual Hallucination Mitigation
Tiantian Dang, Chao Bi, Shufan Shen +3
Despite the significant advancements in Large Vision-Language Models (LVLMs), their tendency to generate hallucinations undermines reliability and restricts broader practical deplo…