5 papers
The Curse of Multiple Mediators: Hidden Interaction Effects in Activation Patching
Sankaran Vaidyanathan, David Arbour, Aaron Mueller +2
Activation patching is the primary tool in mechanistic interpretability. It attributes causal responsibility for a model behavior to each of its individual components by estimating…
From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?
Aaron Mueller, Andrew Lee, Shruti Joshi +3
A goal of interpretability is to recover disentangled representations of latent concepts (features) from the activations of neural networks. The quality of features is typically ev…
Do Language Models Track Entities Across State Changes?
Zilu Tang, Qiao Zhao, Gabriel Franco +4
Entity tracking (ET), the ability to keep track of states, is a fundamental skill that underlies complex reasoning. An increasing amount of work investigates how transformer langua…
Latent Agents: A Post-Training Procedure for Internalized Multi-Agent Debate
John Seon Keun Yi, Aaron Mueller, Dokyun Lee
Multi-agent debate has been shown to improve reasoning in large language models (LLMs). However, it is compute-intensive, requiring generation of long transcripts before answering…
Priors in Time: Missing Inductive Biases for Language Model Interpretability
Ekdeep Singh Lubana, Can Rager, Sai Sumedh R. Hindupur +13
Recovering meaningful concepts from language model activations is a central aim of interpretability. While existing feature extraction methods aim to identify concepts that are ind…