1 citations · 1 across the 5 of their papers we have counts for
5 papers
Transcoders Beat Sparse Autoencoders for Interpretability
Gonçalo Paulo, Stepan Shabalin, Nora Belrose
Sparse autoencoders (SAEs) extract human-interpretable features from deep neural networks by transforming their activations into a sparse, higher dimensional latent space, and then…
Slowing Learning by Erasing Simple Features
Lucia Quirke, Nora Belrose
Prior work suggests that neural networks tend to learn low-order moments of the data distribution first, before moving on to higher-order correlations. In this work, we derive a no…
Converting MLPs into Polynomials in Closed Form
Nora Belrose, Alice Rigg
Recent work has shown that purely quadratic functions can replace MLPs in transformers with no significant loss in performance, while enabling new methods of interpretability based…
Partially Rewriting a Transformer in Natural Language
Gonçalo Paulo, Nora Belrose
The greatest ambition of mechanistic interpretability is to completely rewrite deep neural networks in a format that is more amenable to human understanding, while preserving their…
Sparse Autoencoders Trained on the Same Data Learn Different Features
Gonçalo Paulo, Nora Belrose
Sparse autoencoders (SAEs) are a useful tool for uncovering human-interpretable features in the activations of large language models (LLMs). While some expect SAEs to find the true…