1 paper · 1 filter
Youssef Zaazou, Mark Thomas
Vision-language models (VLMs), such as CLIP and SigLIP 2, are widely used for image classification, yet their vision encoders remain vulnerable to systematic biases that undermine…