1 paper
Harsha Vardhan Khurdula, Basem Rizk, Indus Khaitan +3
Current benchmarks for evaluating Vision Language Models (VLMs) often fall short in thoroughly assessing model abilities to understand and process complex visual and textual conten…