1 paper
Raul Ortega, José Manuel Gómez-Pérez
Vision-language models (VLMs) have demonstrated strong performance in visual question answering with natural images. However, they continue to struggle with scientific diagrams, wh…