10 citations · 10 across the 3 of their papers we have counts for
3 papers
cs.LG2026
Learning a Generative Meta-Model of LLM Activations
Grace Luo, Jiahai Feng, Trevor Darrell +2
Existing approaches for analyzing neural network activations, such as PCA and sparse autoencoders, rely on strong structural assumptions. Generative models offer an alternative: th…
cs.LG2026
Shaping capabilities with token-level data filtering
Neil Rathi, Alec Radford
Current approaches to reducing undesired capabilities in language models are largely post hoc, and can thus be easily bypassed by adversaries. A natural alternative is to shape cap…
cs.LG2024★ 10 cited
Scaling and evaluating sparse autoencoders
Leo Gao, Tom Dupré la Tour, Henk Tillman +6
Sparse autoencoders provide a promising unsupervised approach for extracting interpretable features from a language model by reconstructing activations from a sparse bottleneck lay…