1 paper · 1 filter
David Chanin, James Wilken-Smith, Tomáš Dulka +3
Sparse Autoencoders (SAEs) aim to decompose the activation space of large language models (LLMs) into human-interpretable latent directions or features. As we increase the number o…