4 papers
CycleCap: Improving VLMs Captioning Performance via Self-Supervised Cycle Consistency Fine-Tuning
Marios Krestenitis, Christos Tzelepis, Konstantinos Ioannidis +5
Visual-Language Models (VLMs) have achieved remarkable progress in image captioning, visual question answering, and visual reasoning. Yet they remain prone to vision-language misal…
Deconstructing the Failure of Ideal Noise Correction: A Three-Pillar Diagnosis
Chen Feng, Zhuo Zhi, Zhao Huang +5
Statistically consistent methods based on the noise transition matrix () offer a theoretically grounded solution to Learning with Noisy Labels (LNL), with guarantees of converge…
Breaking Language Barriers or Reinforcing Bias? A Study of Gender and Racial Disparities in Multilingual Contrastive Vision Language Models
Zahraa Al Sahili, Ioannis Patras, Matthew Purver
Multilingual vision-language models (VLMs) promise universal image-text retrieval, yet their social biases remain underexplored. We perform the first systematic audit of four publi…
Data Matters Most: Auditing Social Bias in Contrastive Vision Language Models
Zahraa Al Sahili, Ioannis Patras, Matthew Purver
Vision-language models (VLMs) deliver strong zero-shot recognition but frequently inherit social biases from their training data. We systematically disentangle three design factors…