5 citations · 19 across the 16 of their papers we have counts for
25 papers
A Simple Data Augmentation Strategy for Text-in-Image Scientific VQA
Belal Shoer, Yova Kementchedjhieva
Scientific visual question answering poses significant challenges for vision-language models due to the complexity of scientific figures and their multimodal context. Traditional a…
Noise is an Efficient Learner for Zero-Shot Vision-Language Models
Raza Imam, Asif Hanif, Jian Zhang +3
Recently, test-time adaptation has garnered attention as a method for tuning models without labeled data. The conventional modus operandi for adapting pre-trained vision-language m…
EFSA: Episodic Few-Shot Adaptation for Text-to-Image Retrieval
Muhammad Huzaifa, Yova Kementchedjhieva
Text-to-image retrieval is a critical task for managing diverse visual content, but common benchmarks for the task rely on small, single-domain datasets that fail to capture real-w…
CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark
David Romero, Chenyang Lyu, Haryo Akbarianto Wibowo +73
Visual Question Answering (VQA) is an important task in multimodal AI, and it is often used to test the ability of vision-language models to understand and reason on knowledge pres…
Multimodal Large Language Models to Support Real-World Fact-Checking
Jiahui Geng, Yova Kementchedjhieva, Preslav Nakov +1
Multimodal large language models (MLLMs) carry the potential to support humans in processing vast amounts of information. While MLLMs are already being used as a fact-checking tool…
MuLan: A Study of Fact Mutability in Language Models
Constanza Fierro, Nicolas Garneau, Emanuele Bugliarello +2
Facts are subject to contingencies and can be true or false in different circumstances. One such contingency is time, wherein some facts mutate over a given period, e.g., the presi…