papers

Publications (75)

cs.LG2024

Interpreting Attention Layer Outputs with Sparse Autoencoders

Connor Kissane, Robert Krzyzanowski, Joseph Isaac Bloom +2

Decomposing model activations into interpretable components is a key open problem in mechanistic interpretability. Sparse autoencoders (SAEs) are a popular method for decomposing t…

cs.LG2026

Building Production-Ready Probes For Gemini

János Kramár, Joshua Engels, Zheng Wang +4

Frontier language model capabilities are improving rapidly. We thus need stronger mitigations against bad actors misusing increasingly powerful systems. Prior work has shown that a…

cs.LG2024

Towards Best Practices of Activation Patching in Language Models: Metrics and Methods

Fred Zhang, Neel Nanda

Mechanistic interpretability seeks to understand the internal mechanisms of machine learning models, where localization -- identifying the important model components -- is a key st…

cs.LG2024

Confidence Regulation Neurons in Language Models

Alessandro Stolfo, Ben Wu, Wes Gurnee +4

Despite their widespread use, the mechanisms by which large language models (LLMs) represent and regulate uncertainty in next-token predictions remain largely unexplored. This stud…

cs.LG2024

Refusal in Language Models Is Mediated by a Single Direction

Andy Arditi, Oscar Obeso, Aaquib Syed +4

Conversational large language models are fine-tuned for both instruction-following and safety, resulting in models that obey benign requests but refuse harmful ones. While this ref…

cs.LG2024

Summing Up the Facts: Additive Mechanisms Behind Factual Recall in LLMs

Bilal Chughtai, Alan Cooney, Neel Nanda

How do transformer-based large language models (LLMs) store and retrieve knowledge? We focus on the most basic form of this task -- factual recall, where the model is tasked with e…