11 papers
Multimodal Model Diffing for Feature Discovery and Control
Hunar Batra, Lachin Naghashyar, Ashkan Khakzar +4
Multimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause these behaviors remain difficult to identify, audit, or control.…
Base Models Know How to Reason, Thinking Models Learn When
Constantin Venhoff, Iván Arcuschin, Philip Torr +2
What do thinking language models learn during training that their base models lack? We first present an unsupervised method that discovers a model's reasoning behaviors by training…
Probing the Misaligned Thinking Process of Language Models
Kaiwen Zhou, Constantin Venhoff, Jonathan Michala +2
Large language models exhibit a growing range of misaligned behaviors such as strategic deception, sandbagging, and self-preservation. As they are increasingly deployed in high-sta…
Dyslexify: A Mechanistic Defense Against Typographic Attacks in CLIP
Lorenz Hufe, Constantin Venhoff, Erblina Purelku +3
Typographic attacks exploit multi-modal systems by injecting text into images, leading to targeted misclassifications, malicious content generation and even Vision-Language Model j…
Towards Understanding Multimodal Fine-Tuning: Spatial Features
Lachin Naghashyar, Hunar Batra, Ashkan Khakzar +4
Contemporary Vision-Language Models (VLMs) achieve strong performance on a wide range of tasks by pairing a vision encoder with a pre-trained language model, fine-tuned for visual-…
Sparse CLIP: Co-Optimizing Interpretability and Performance in Contrastive Learning
Chuan Qin, Constantin Venhoff, Sonia Joseph +2
Contrastive Language-Image Pre-training (CLIP) has become a cornerstone in vision-language representation learning, powering diverse downstream tasks and serving as the default vis…