1 citations · 1 across the 5 of their papers we have counts for
4 papers · 1 filter
Automated Interpretability Metrics Do Not Distinguish Trained and Random Transformers
Thomas Heap, Tim Lawson, Lucy Farnik +1
Sparse autoencoders (SAEs) are widely used to extract sparse, interpretable latents from transformer activations. We test whether commonly used SAE quality metrics and automatic ex…
Learning to Skip the Middle Layers of Transformers
Tim Lawson, Laurence Aitchison
Conditional computation is a popular strategy to make Transformers more efficient. Existing methods often target individual modules (e.g., mixture-of-experts layers) or skip layers…
Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations
Lucy Farnik, Tim Lawson, Conor Houghton +1
Sparse autoencoders (SAEs) have been successfully used to discover sparse and human-interpretable representations of the latent activations of LLMs. However, we would ultimately li…
Residual Stream Analysis with Multi-Layer SAEs
Tim Lawson, Lucy Farnik, Conor Houghton +1
Sparse autoencoders (SAEs) are a promising approach to interpreting the internal representations of transformer language models. However, SAEs are usually trained separately on eac…