Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
Unexplored flaws in multiple-choice VQA make benchmarking unreliable
Fabio Rosenthal, Sebastian Schmidt, Thorsten Graf +3
Previous works identify sensitivity to option order as a key issue in multiple-choice VQA (MC-VQA) evaluation and propose protocols to mitigate this effect. We show that such mitig…
cs.CV2025
FOCUS: Internal MLLM Representations for Efficient Fine-Grained Visual Question Answering
Liangyu Zhong, Fabio Rosenthal, Joachim Sicking +4
While Multimodal Large Language Models (MLLMs) offer strong perception and reasoning capabilities for image-text input, Visual Question Answering (VQA) focusing on small image deta…