3 citations · 3 across the 1 of their papers we have counts for
1 paper · 1 filter
Kevin Der, Harish Kamath, Ben Thompson
Sparse autoencoders (SAEs) have become a useful tool for extracting interpretable features in language models. However, standard SAE architectures operate on individual token activ…