6 papers
Ensembling Sparse Autoencoders
Soham Gadgil, Chris Lin, Su-In Lee
Sparse autoencoders (SAEs) are used to decompose neural network activations into human-interpretable features. Typically, features learned by a single SAE are used for downstream a…
A Geometric View for Understanding Concept Learning and Neuron Interpretation in Sparse Autoencoders
Chenhao Zhang, Chris Lin, Su-In Lee
We propose a unified mathematical framework for a geometric understanding of concept learning and neuron interpretation in sparse autoencoders (SAEs). While SAEs improve interpreta…
SurrogateSHAP: Training-Free Contributor Attribution for Text-to-Image (T2I) Models
Mingyu Lu, Soham Gadgil, Chris Lin +2
As Text-to-Image (T2I) diffusion models are increasingly used in real-world creative workflows, a principled framework for valuing contributors who provide a collection of data is…
Agents that Matter: Optimizing Multi-Agent LLMs via Removal-Based Attribution
Mingyu Lu, Yushan Huang, Chris Lin +1
As multi-agent systems (MAS) become increasingly complex, identifying the contributions of individual agents is critical for system optimization. However, existing approaches lack…
Where to Steer: Input-Dependent Layer Selection for Steering Improves LLM Alignment
Soham Gadgil, Chris Lin, Su-In Lee
Steering vectors have emerged as a lightweight and effective approach for aligning large language models (LLMs) at inference time, enabling modulation over model behaviors by shift…
An Efficient Framework for Crediting Data Contributors of Diffusion Models
Chris Lin, Mingyu Lu, Chanwoo Kim +1
As diffusion models are deployed in real-world settings, and their performance is driven by training data, appraising the contribution of data contributors is crucial to creating i…