1 paper · 1 filter
Naren Kumar S, Tirth Bhatt, Mayank Singh
Vision-language models can describe an image with remarkable accuracy, yet a more fundamental question remains unanswered: what visual information actually drives their answers? In…