2 papers
cs.CV2025
Discovering Meaningful Units with Visually Grounded Semantics from Image Captions
Melika Behjati, James Henderson
Fine-grained knowledge is crucial for vision-language models to obtain a better understanding of the real world. While there has been work trying to acquire this kind of knowledge…
cs.CV2024
GABInsight: Exploring Gender-Activity Binding Bias in Vision-Language Models
Ali Abdollahi, Mahdi Ghaznavi, Mohammad Reza Karimi Nejad +6
Vision-language models (VLMs) are intensively used in many downstream tasks, including those requiring assessments of individuals appearing in the images. While VLMs perform well i…