1 paper · 1 filter
Tao Zhang, Ziqi Zhang, Zongyang Ma +10
Advanced Multimodal Large Language Models (MLLMs) struggle with recent Knowledge-based Visual Question Answering (VQA) tasks, such as INFOSEEK and Encyclopedic-VQA, due to their li…