3 papers
cs.CV2025
EFSA: Episodic Few-Shot Adaptation for Text-to-Image Retrieval
Muhammad Huzaifa, Yova Kementchedjhieva
Text-to-image retrieval is a critical task for managing diverse visual content, but common benchmarks for the task rely on small, single-domain datasets that fail to capture real-w…
cs.CV2025
A Simple Data Augmentation Strategy for Text-in-Image Scientific VQA
Belal Shoer, Yova Kementchedjhieva
Scientific visual question answering poses significant challenges for vision-language models due to the complexity of scientific figures and their multimodal context. Traditional a…
cs.CV2025
Noise is an Efficient Learner for Zero-Shot Vision-Language Models
Raza Imam, Asif Hanif, Jian Zhang +3
Recently, test-time adaptation has garnered attention as a method for tuning models without labeled data. The conventional modus operandi for adapting pre-trained vision-language m…