1 paper · 1 filter
Ahmed Oumar El-Shangiti, Abzal Nurgazy, Hilal AlQuabeh +2
Despite strong performance on many multimodal tasks, vision-language models (VLMs) still struggle with basic object counting. We investigate whether this reflects missing internal…