1 paper
Liangyu Zhong, Fabio Rosenthal, Joachim Sicking +4
While Multimodal Large Language Models (MLLMs) offer strong perception and reasoning capabilities for image-text input, Visual Question Answering (VQA) focusing on small image deta…