4 papers · 1 filter
Binary Sparse Coding for Interpretability
Lucia Quirke, Stepan Shabalin, Nora Belrose
Sparse autoencoders (SAEs) are used to decompose neural network activations into sparsely activating features, but many SAE features are only interpretable at high activation stren…
Interpreting Large Text-to-Image Diffusion Models with Dictionary Learning
Stepan Shabalin, Ayush Panda, Dmitrii Kharlapenko +3
Sparse autoencoders are a promising new approach for decomposing language model activations for interpretation and control. They have been applied successfully to vision transforme…
Scaling sparse feature circuit finding for in-context learning
Dmitrii Kharlapenko, Stepan Shabalin, Fazl Barez +2
Sparse autoencoders (SAEs) are a popular tool for interpreting large language model activations, but their utility in addressing open questions in interpretability remains unclear.…
Transcoders Beat Sparse Autoencoders for Interpretability
Gonçalo Paulo, Stepan Shabalin, Nora Belrose
Sparse autoencoders (SAEs) extract human-interpretable features from deep neural networks by transforming their activations into a sparse, higher dimensional latent space, and then…