4 papers
Unstable Features, Reproducible Subspaces: Understanding Seed Dependence in Sparse Autoencoders
Gleb Gerasimov, Timofei Rusalev, Nikita Balagansky +3
Sparse autoencoders (SAEs) are widely used to interpret neural network representations, but their utility depends on whether the learned features are reproducible across training r…
Small Vectors, Big Effects: A Mechanistic Study of RL-Induced Reasoning via Steering Vectors
Viacheslav Sinii, Nikita Balagansky, Gleb Gerasimov +6
The mechanisms by which reasoning training reshapes LLMs' internal computations remain unclear. We study lightweight steering vectors inserted into the base model's residual stream…
Train One Sparse Autoencoder Across Multiple Sparsity Budgets to Preserve Interpretability and Accuracy
Nikita Balagansky, Yaroslav Aksenov, Daniil Laptev +4
Sparse Autoencoders (SAEs) have proven to be powerful tools for interpreting neural networks by decomposing hidden representations into disentangled, interpretable features via spa…
Kronecker Factorization Improves Efficiency and Interpretability of Sparse Autoencoders
Vadim Kurochkin, Yaroslav Aksenov, Daniil Laptev +2
Sparse Autoencoders (SAEs) decompose language-model activations into sparse, interpretable features, but standard encoders usually treat the latent dictionary as a flat set of inde…