5 papers · 1 filter
Ensembling Sparse Autoencoders
Soham Gadgil, Chris Lin, Su-In Lee
Sparse autoencoders (SAEs) are used to decompose neural network activations into human-interpretable features. Typically, features learned by a single SAE are used for downstream a…
A Geometric View for Understanding Concept Learning and Neuron Interpretation in Sparse Autoencoders
Chenhao Zhang, Chris Lin, Su-In Lee
We propose a unified mathematical framework for a geometric understanding of concept learning and neuron interpretation in sparse autoencoders (SAEs). While SAEs improve interpreta…
SurrogateSHAP: Training-Free Contributor Attribution for Text-to-Image (T2I) Models
Mingyu Lu, Soham Gadgil, Chris Lin +2
As Text-to-Image (T2I) diffusion models are increasingly used in real-world creative workflows, a principled framework for valuing contributors who provide a collection of data is…
Where to Steer: Input-Dependent Layer Selection for Steering Improves LLM Alignment
Soham Gadgil, Chris Lin, Su-In Lee
Steering vectors have emerged as a lightweight and effective approach for aligning large language models (LLMs) at inference time, enabling modulation over model behaviors by shift…
An Efficient Framework for Crediting Data Contributors of Diffusion Models
Chris Lin, Mingyu Lu, Chanwoo Kim +1
As diffusion models are deployed in real-world settings, and their performance is driven by training data, appraising the contribution of data contributors is crucial to creating i…