3 citations · 3 across the 4 of their papers we have counts for
1 paper · 2 filters
Kevin Der, Harish Kamath, Ben Thompson
Sparse autoencoders (SAEs) have become a useful tool for extracting interpretable features in language models. However, standard SAE architectures operate on individual token activ…