1 paper
Duy Le Dinh Anh, Patrick Amadeus Irawan, Tuan Van Vo
Vision--language models (VLMs) have achieved impressive performance on complex multimodal reasoning tasks, yet they still fail on simple grounding skills such as object counting. E…