1 paper · 1 filter
Kung-Hsiang Huang, Can Qin, Haoyi Qiu +4
Vision Language Models (VLMs) have achieved remarkable progress in multimodal tasks, yet they often struggle with visual arithmetic, seemingly simple capabilities like object count…