1 paper · 1 filter
Boyuan Chen, Minghao Shao, Siddharth Garg +2
Vision Language Models (VLMs) exhibit persistent hallucinations in counting tasks, with accuracy substantially lower than other visual reasoning tasks (excluding sentiment). This p…