3 papers
cs.AI2025
Interpretable Embeddings with Sparse Autoencoders: A Data Analysis Toolkit
Nick Jiang, Xiaoqing Sun, Lisa Dunlap +2
Analyzing large-scale text corpora is a core challenge in machine learning, crucial for tasks like identifying undesirable model behaviors or biases in training data. Current metho…
cs.LG2025
Dense SAE Latents Are Features, Not Bugs
Xiaoqing Sun, Alessandro Stolfo, Joshua Engels +4
Sparse autoencoders (SAEs) are designed to extract interpretable features from language models by enforcing a sparsity constraint. Ideally, training an SAE would yield latents that…
q-bio.NC2025
The Geometry of Concepts: Sparse Autoencoder Feature Structure
Yuxiao Li, Eric J. Michaud, David D. Baek +3
Sparse autoencoders have recently produced dictionaries of high-dimensional vectors corresponding to the universe of concepts represented by large language models. We find that thi…