1 citations · 1 across the 3 of their papers we have counts for
1 paper · 1 filter
Tao Zhang, Ziqi Zhang, Zongyang Ma +10
Advanced Multimodal Large Language Models (MLLMs) struggle with recent Knowledge-based Visual Question Answering (VQA) tasks, such as INFOSEEK and Encyclopedic-VQA, due to their li…