4 papers · 1 filter
ROSA: Addressing text understanding challenges in photographs via ROtated SAmpling
Hernán Maina, Guido Ivetta, Mateo Lione Stuto +3
Visually impaired people could benefit from Visual Question Answering (VQA) systems to interpret text in their surroundings. However, current models often struggle with recognizing…
HESEIA: A community-based dataset for evaluating social biases in large language models, co-designed in real school settings in Latin America
Guido Ivetta, Marcos J. Gomez, Sofía Martinelli +5
Most resources for evaluating social biases in Large Language Models are developed without co-design from the communities affected by these biases, and rarely involve participatory…
CaMMT: Benchmarking Culturally Aware Multimodal Machine Translation
Emilio Villa-Cueva, Sholpan Bolatzhanova, Diana Turmakhan +32
Translating cultural content poses challenges for machine translation systems due to the differences in conceptualizations between cultures, where language alone may fail to convey…
Selectively Answering Visual Questions
Julian Martin Eisenschlos, Hernán Maina, Guido Ivetta +1
Recently, large multi-modal models (LMMs) have emerged with the capacity to perform vision tasks such as captioning and visual question answering (VQA) with unprecedented accuracy.…