2 papers
cs.LG2026
A Geometric View for Understanding Concept Learning and Neuron Interpretation in Sparse Autoencoders
Chenhao Zhang, Chris Lin, Su-In Lee
We propose a unified mathematical framework for a geometric understanding of concept learning and neuron interpretation in sparse autoencoders (SAEs). While SAEs improve interpreta…
cs.LG2024
Stochastic Amortization: A Unified Approach to Accelerate Feature and Data Attribution
Ian Covert, Chanwoo Kim, Su-In Lee +2
Many tasks in explainable machine learning, such as data valuation and feature attribution, perform expensive computation for each data point and are intractable for large datasets…