10 citations · 10 across the 3 of their papers we have counts for
1 paper · 1 filter
Wilson E. Marcílio-Jr, Danilo M. Eler
Sparse autoencoders (SAEs) trained on large language model activations output thousands of features that enable mapping to human-interpretable concepts. The current practice for an…