◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Stepan Shabalin

6 papers hereh-index 5322 citations9 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author1
  • middle author5

Across the 6 of 6 papers where every author was matched, so the position is known.

fields
  • cs.LG4
  • cs.AI1
  • cs.CL1

identity via Semantic Scholar / OpenAlex

collaborators
Showing cs.LGShow all

4 papers · 1 filter

cs.LG2025

Binary Sparse Coding for Interpretability

Lucia Quirke, Stepan Shabalin, Nora Belrose

Sparse autoencoders (SAEs) are used to decompose neural network activations into sparsely activating features, but many SAE features are only interpretable at high activation stren…

cs.LG2025

Interpreting Large Text-to-Image Diffusion Models with Dictionary Learning

Stepan Shabalin, Ayush Panda, Dmitrii Kharlapenko +3

Sparse autoencoders are a promising new approach for decomposing language model activations for interpretation and control. They have been applied successfully to vision transforme…

cs.LG2025

Scaling sparse feature circuit finding for in-context learning

Dmitrii Kharlapenko, Stepan Shabalin, Fazl Barez +2

Sparse autoencoders (SAEs) are a popular tool for interpreting large language model activations, but their utility in addressing open questions in interpretability remains unclear.…

cs.LG2025

Transcoders Beat Sparse Autoencoders for Interpretability

Gonçalo Paulo, Stepan Shabalin, Nora Belrose

Sparse autoencoders (SAEs) extract human-interpretable features from deep neural networks by transforming their activations into a sparse, higher dimensional latent space, and then…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.