5 papers
Binary Sparse Coding for Interpretability
Lucia Quirke, Stepan Shabalin, Nora Belrose
Sparse autoencoders (SAEs) are used to decompose neural network activations into sparsely activating features, but many SAE features are only interpretable at high activation stren…
Interpreting Large Text-to-Image Diffusion Models with Dictionary Learning
Stepan Shabalin, Ayush Panda, Dmitrii Kharlapenko +3
Sparse autoencoders are a promising new approach for decomposing language model activations for interpretation and control. They have been applied successfully to vision transforme…
Patterns and Mechanisms of Contrastive Activation Engineering
Yixiong Hao, Ayush Panda, Stepan Shabalin +1
Controlling the behavior of Large Language Models (LLMs) remains a significant challenge due to their inherent complexity and opacity. While techniques like fine-tuning can modify…
Scaling sparse feature circuit finding for in-context learning
Dmitrii Kharlapenko, Stepan Shabalin, Fazl Barez +2
Sparse autoencoders (SAEs) are a popular tool for interpreting large language model activations, but their utility in addressing open questions in interpretability remains unclear.…
Transcoders Beat Sparse Autoencoders for Interpretability
Gonçalo Paulo, Stepan Shabalin, Nora Belrose
Sparse autoencoders (SAEs) extract human-interpretable features from deep neural networks by transforming their activations into a sparse, higher dimensional latent space, and then…