activity
20242026
collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

Multimodal Model Diffing for Feature Discovery and Control

Hunar Batra, Lachin Naghashyar, Ashkan Khakzar +4

Multimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause these behaviors remain difficult to identify, audit, or control.…

cs.CV2026

Dyslexify: A Mechanistic Defense Against Typographic Attacks in CLIP

Lorenz Hufe, Constantin Venhoff, Erblina Purelku +3

Typographic attacks exploit multi-modal systems by injecting text into images, leading to targeted misclassifications, malicious content generation and even Vision-Language Model j…

cs.CV2026

Towards Understanding Multimodal Fine-Tuning: Spatial Features

Lachin Naghashyar, Hunar Batra, Ashkan Khakzar +4

Contemporary Vision-Language Models (VLMs) achieve strong performance on a wide range of tasks by pairing a vision encoder with a pre-trained language model, fine-tuned for visual-…

cs.CV2026

Sparse CLIP: Co-Optimizing Interpretability and Performance in Contrastive Learning

Chuan Qin, Constantin Venhoff, Sonia Joseph +2

Contrastive Language-Image Pre-training (CLIP) has become a cornerstone in vision-language representation learning, powering diverse downstream tasks and serving as the default vis…

cs.CV2025

How Visual Representations Map to Language Feature Space in Multimodal LLMs

Constantin Venhoff, Ashkan Khakzar, Sonia Joseph +2

Effective multimodal reasoning depends on the alignment of visual and linguistic representations, yet the mechanisms by which vision-language models (VLMs) achieve this alignment r…