1 paper · 1 filter
Claire Tian, Katherine Tian, Nathan Hu
Sparse Autoencoder (SAE) features have become essential tools for mechanistic interpretability research. SAE features are typically characterized by examining their activating exam…