1 paper · 1 filter
Simon Park, Abhishek Panigrahi, Yun Cheng +3
Vision Language Models (VLMs) are impressive at visual question answering and image captioning. But they underperform on multi-step visual reasoning -- even compared to LLMs on the…