2 citations · 2 across the 4 of their papers we have counts for
5 papers · 1 filter
Priors in Time: Missing Inductive Biases for Language Model Interpretability
Ekdeep Singh Lubana, Can Rager, Sai Sumedh R. Hindupur +13
Recovering meaningful concepts from language model activations is a central aim of interpretability. While existing feature extraction methods aim to identify concepts that are ind…
Automatically Finding Rule-Based Neurons in OthelloGPT
Aditya Singh, Zihang Wen, Srujananjali Medicherla +2
OthelloGPT, a transformer trained to predict valid moves in Othello, provides an ideal testbed for interpretability research. The model is complex enough to exhibit rich computatio…
SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability
Adam Karvonen, Can Rager, Johnny Lin +12
Sparse autoencoders (SAEs) are a popular technique for interpreting language model activations, and there is extensive recent work on improving SAE effectiveness. However, most pri…
Evaluating Sparse Autoencoders on Targeted Concept Erasure Tasks
Adam Karvonen, Can Rager, Samuel Marks +1
Sparse Autoencoders (SAEs) are an interpretability technique aimed at decomposing neural network activations into interpretable units. However, a major bottleneck for SAE developme…
Attribution Patching Outperforms Automated Circuit Discovery
Aaquib Syed, Can Rager, Arthur Conmy
Automated interpretability research has recently attracted attention as a potential research direction that could scale explanations of neural network behavior to large models. Exi…