11 papers
Unstable Features, Reproducible Subspaces: Understanding Seed Dependence in Sparse Autoencoders
Gleb Gerasimov, Timofei Rusalev, Nikita Balagansky +3
Sparse autoencoders (SAEs) are widely used to interpret neural network representations, but their utility depends on whether the learned features are reproducible across training r…
Interpreting and Steering a Text-to-Speech Language Model with Sparse Autoencoders
Nikita Koriagin, Georgii Aparin, Nikita Balagansky +1
Language models increasingly serve as the backbone of text-to-speech (TTS) systems, yet we understand little about the representations they build when text and generated speech tok…
Guided Star-Shaped Masked Diffusion
Viacheslav Meshchaninov, Egor Shibaev, Artem Makoian +5
The performance of pre-trained masked diffusion models is often constrained by their sampling procedure, which makes decisions irreversible and struggles in low-step generation reg…
Kronecker Factorization Improves Efficiency and Interpretability of Sparse Autoencoders
Vadim Kurochkin, Yaroslav Aksenov, Daniil Laptev +2
Sparse Autoencoders (SAEs) have demonstrated significant promise in interpreting the hidden states of language models by decomposing them into interpretable latent directions. Howe…
Steering LLM Reasoning Through Bias-Only Adaptation
Viacheslav Sinii, Alexey Gorbatovski, Artem Cherepanov +3
We show that training a single -dimensional steering vector per layer with reinforcement learning, while freezing all base weights, matches the accuracy of fully RL-tuned reason…
Analyze Feature Flow to Enhance Interpretation and Steering in Language Models
Daniil Laptev, Nikita Balagansky, Yaroslav Aksenov +1
We introduce a new approach to systematically map features discovered by sparse autoencoder across consecutive layers of large language models, extending earlier work that examined…