activity
20242026
collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG2025

Too Late to Recall: Explaining the Two-Hop Problem in Multimodal Knowledge Retrieval

Constantin Venhoff, Ashkan Khakzar, Sonia Joseph +2

Training vision language models (VLMs) aims to align visual representations from a vision encoder with the textual representations of a pretrained large language model (LLM). Howev…

cs.LG2025

Understanding Reasoning in Thinking Language Models via Steering Vectors

Constantin Venhoff, Iván Arcuschin, Philip Torr +2

Recent advances in large language models (LLMs) have led to the development of thinking language models that generate extensive internal reasoning chains before producing responses…

cs.LG2025

Reasoning-Finetuning Repurposes Latent Representations in Base Models

Jake Ward, Chuqiao Lin, Constantin Venhoff +1

Backtracking, an emergent behavior elicited by reasoning fine-tuning, has been shown to be a key mechanism in reasoning models' enhanced capabilities. Prior work has succeeded in m…

cs.LG2025

Mixture of Experts Made Intrinsically Interpretable

Xingyi Yang, Constantin Venhoff, Ashkan Khakzar +4

Neurons in large language models often exhibit \emph{polysemanticity}, simultaneously encoding multiple unrelated concepts and obscuring interpretability. Instead of relying on pos…

cs.LG2024

SAGE: Scalable Ground Truth Evaluations for Large Sparse Autoencoders

Constantin Venhoff, Anisoara Calinescu, Philip Torr +1

A key challenge in interpretability is to decompose model activations into meaningful features. Sparse autoencoders (SAEs) have emerged as a promising tool for this task. However,…