1 paper · 1 filter
Shalom Kachko, Raz Lapid, Margarita Vald +2
Vision-language models (VLMs) process image patches and text tokens in a shared residual stream, but the local geometry through which the two modalities interact remains poorly und…