1 paper
Ombretta Strafforello, Derya Soydaner, Michiel Willems +2
The emergence of large Vision-Language Models (VLMs) has established new baselines in image classification across multiple domains. We examine whether their multimodal reasoning ca…