3 papers
cs.CV2026
Sparse CLIP: Co-Optimizing Interpretability and Performance in Contrastive Learning
Chuan Qin, Constantin Venhoff, Sonia Joseph +2
Contrastive Language-Image Pre-training (CLIP) has become a cornerstone in vision-language representation learning, powering diverse downstream tasks and serving as the default vis…
cs.LG2025
Too Late to Recall: Explaining the Two-Hop Problem in Multimodal Knowledge Retrieval
Constantin Venhoff, Ashkan Khakzar, Sonia Joseph +2
Training vision language models (VLMs) aims to align visual representations from a vision encoder with the textual representations of a pretrained large language model (LLM). Howev…
cs.CV2025
How Visual Representations Map to Language Feature Space in Multimodal LLMs
Constantin Venhoff, Ashkan Khakzar, Sonia Joseph +2
Effective multimodal reasoning depends on the alignment of visual and linguistic representations, yet the mechanisms by which vision-language models (VLMs) achieve this alignment r…