most citedBecause we have LLMs, we Can and Should Pursue Agentic Interpretability

1 citations · 1 across the 3 of their papers we have counts for

collaborators

7 papers

cs.LG2025

Too Late to Recall: Explaining the Two-Hop Problem in Multimodal Knowledge Retrieval

Constantin Venhoff, Ashkan Khakzar, Sonia Joseph +2

Training vision language models (VLMs) aims to align visual representations from a vision encoder with the textual representations of a pretrained large language model (LLM). Howev…

cs.AI2025

Internal states before wait modulate reasoning patterns

Dmitrii Troitskii, Koyena Pal, Chris Wendler +2

Prior work has shown that a significant driver of performance in reasoning models is their ability to reason and self-correct. A distinctive marker in these reasoning traces is the…

cs.LG2025

RelP: Faithful and Efficient Circuit Discovery in Language Models via Relevance Patching

Farnoush Rezaei Jafari, Oliver Eberle, Ashkan Khakzar +1

Activation patching is a standard method in mechanistic interpretability for localizing the components of a model responsible for specific behaviors, but it is computationally expe…

cs.LG2025

Reasoning-Finetuning Repurposes Latent Representations in Base Models

Jake Ward, Chuqiao Lin, Constantin Venhoff +1

Backtracking, an emergent behavior elicited by reasoning fine-tuning, has been shown to be a key mechanism in reasoning models' enhanced capabilities. Prior work has succeeded in m…

cs.CL2025

Simple Mechanistic Explanations for Out-Of-Context Reasoning

Atticus Wang, Joshua Engels, Oliver Clive-Griffin +2

Out-of-context reasoning (OOCR) is a phenomenon in which fine-tuned LLMs exhibit surprisingly deep out-of-distribution generalization. Rather than learning shallow heuristics, they…

cs.CV2025

How Visual Representations Map to Language Feature Space in Multimodal LLMs

Constantin Venhoff, Ashkan Khakzar, Sonia Joseph +2

Effective multimodal reasoning depends on the alignment of visual and linguistic representations, yet the mechanisms by which vision-language models (VLMs) achieve this alignment r…