1 paper · 1 filter
Kairui Zhang, Ziwen Yu, Zahraa S. Abdallah +1
Sparse autoencoders (SAEs) provide useful decompositions of Transformer residual streams, but their learned features are usually named post hoc rather than directly connected to th…