collaborators

5 papers

cs.LG2025

Binary Sparse Coding for Interpretability

Lucia Quirke, Stepan Shabalin, Nora Belrose

Sparse autoencoders (SAEs) are used to decompose neural network activations into sparsely activating features, but many SAE features are only interpretable at high activation stren…

cs.LG2025

Interpreting Large Text-to-Image Diffusion Models with Dictionary Learning

Stepan Shabalin, Ayush Panda, Dmitrii Kharlapenko +3

Sparse autoencoders are a promising new approach for decomposing language model activations for interpretation and control. They have been applied successfully to vision transforme…

cs.AI2025

Patterns and Mechanisms of Contrastive Activation Engineering

Yixiong Hao, Ayush Panda, Stepan Shabalin +1

Controlling the behavior of Large Language Models (LLMs) remains a significant challenge due to their inherent complexity and opacity. While techniques like fine-tuning can modify…

cs.LG2025

Scaling sparse feature circuit finding for in-context learning

Dmitrii Kharlapenko, Stepan Shabalin, Fazl Barez +2

Sparse autoencoders (SAEs) are a popular tool for interpreting large language model activations, but their utility in addressing open questions in interpretability remains unclear.…

cs.LG2025

Transcoders Beat Sparse Autoencoders for Interpretability

Gonçalo Paulo, Stepan Shabalin, Nora Belrose

Sparse autoencoders (SAEs) extract human-interpretable features from deep neural networks by transforming their activations into a sparse, higher dimensional latent space, and then…