1 paper
Onat Ozdemir, Anders Christensen, Stephan Alaniz +2
Large-scale vision-language models such as CLIP have achieved remarkable success in zero-shot image recognition, yet their predictions remain largely opaque to human understanding.…