37 citations · 40 across the 11 of their papers we have counts for
1 paper · 1 filter
Ananya Joshi, Celia Cintas, Skyler Speakman
Recent work shows that Sparse Autoencoders (SAE) applied to large language model (LLM) layers have neurons corresponding to interpretable concepts. These SAE neurons can be modifie…