2 papers
cs.CV2026
Gaze Heads: How VLMs Look at What They Describe
Rohit Gandikota, David Bau
How a vision-language model internally solves the task of describing an image is far from obvious. We find that the model develops a specific mechanism for this: a small set of att…
cs.CV2026
The Dual Mechanisms of Spatial Variable Binding in Vision-Language Models
Kelly Cui, Nikhil Prakash, Shoval Messica +4
Many multimodal tasks, such as image captioning and visual question answering, require vision-language models (VLMs) to bind objects with their properties and spatial relations. Ye…