5 papers · 1 filter
A Geometric View for Understanding Concept Learning and Neuron Interpretation in Sparse Autoencoders
Chenhao Zhang, Chris Lin, Su-In Lee
We propose a unified mathematical framework for a geometric understanding of concept learning and neuron interpretation in sparse autoencoders (SAEs). While SAEs improve interpreta…
Where to Steer: Input-Dependent Layer Selection for Steering Improves LLM Alignment
Soham Gadgil, Chris Lin, Su-In Lee
Steering vectors have emerged as a lightweight and effective approach for aligning large language models (LLMs) at inference time, enabling modulation over model behaviors by shift…
SurrogateSHAP: Training-Free Contributor Attribution for Text-to-Image (T2I) Models
Mingyu Lu, Soham Gadgil, Chris Lin +2
As Text-to-Image (T2I) diffusion models are increasingly used in real-world creative workflows, a principled framework for valuing contributors who provide a collection of data is…
Ensembling Sparse Autoencoders
Soham Gadgil, Chris Lin, Su-In Lee
Sparse autoencoders (SAEs) are used to decompose neural network activations into human-interpretable features. Typically, features learned by a single SAE are used for downstream a…
An Efficient Framework for Crediting Data Contributors of Diffusion Models
Chris Lin, Mingyu Lu, Chanwoo Kim +1
As diffusion models are deployed in real-world settings, and their performance is driven by training data, appraising the contribution of data contributors is crucial to creating i…