4 papers
Low-resource domain adaptation while minimizing energy and hardware resource consumption
Hernán Maina, Nicolás Wolovick, Luciana Benotti
Training Large Language Models (LLMs) is costly in terms of energy, hardware, and annotated data, often resulting in a positionality rooted in predominant cultures and values (Sant…
ROSA: Addressing text understanding challenges in photographs via ROtated SAmpling
Hernán Maina, Guido Ivetta, Mateo Lione Stuto +3
Visually impaired people could benefit from Visual Question Answering (VQA) systems to interpret text in their surroundings. However, current models often struggle with recognizing…
Selectively Answering Visual Questions
Julian Martin Eisenschlos, Hernán Maina, Guido Ivetta +1
Recently, large multi-modal models (LMMs) have emerged with the capacity to perform vision tasks such as captioning and visual question answering (VQA) with unprecedented accuracy.…
CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark
David Romero, Chenyang Lyu, Haryo Akbarianto Wibowo +73
Visual Question Answering (VQA) is an important task in multimodal AI, and it is often used to test the ability of vision-language models to understand and reason on knowledge pres…