12 citations · 48 across the 58 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
State-Conditioned Visual Evidence Retrieval for Fine-Grained Perception in Document Vision-Language Models
Mingxu Chai, Chenyu Liu, Ziyu Shen +7
Compared with typical vision-language tasks, document parsing places stronger demands on fine-grained visual perception. Existing vision-language model (VLM)-based parsing approach…
cs.CV2025
Counteracting Matthew Effect in Self-Improvement of LVLMs through Head-Tail Re-balancing
Xin Guo, Zhiheng Xi, Yiwen Ding +6
Self-improvement has emerged as a mainstream paradigm for advancing the reasoning capabilities of large vision-language models (LVLMs), where models explore and learn from successf…
cs.CV2024
Have the VLMs Lost Confidence? A Study of Sycophancy in VLMs
Shuo Li, Tao Ji, Xiaoran Fan +10
In the study of LLMs, sycophancy represents a prevalent hallucination that poses significant challenges to these models. Specifically, LLMs often fail to adhere to original correct…