1 paper · 1 filter
Grégoire Dhimoïla, Thomas Fel, Victor Boutin +1
Vision-language models (VLMs) align images and text with remarkable success, yet the geometry of their shared embedding space remains poorly understood. To probe this geometry, we…