Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Sparse CLIP: Co-Optimizing Interpretability and Performance in Contrastive Learning
Chuan Qin, Constantin Venhoff, Sonia Joseph +2
Contrastive Language-Image Pre-training (CLIP) has become a cornerstone in vision-language representation learning, powering diverse downstream tasks and serving as the default vis…
cs.CV2025
How Visual Representations Map to Language Feature Space in Multimodal LLMs
Constantin Venhoff, Ashkan Khakzar, Sonia Joseph +2
Effective multimodal reasoning depends on the alignment of visual and linguistic representations, yet the mechanisms by which vision-language models (VLMs) achieve this alignment r…