1 paper · 1 filter
Pratham Singla, Shivank Garg, Vihan Singh +1
Vision-language models (VLMs) are increasingly deployed where answers must follow from what is in the image, yet they often answer from textual priors, the question's phrasing toge…