2 papers
cs.LG2025
Evaluating and Designing Sparse Autoencoders by Approximating Quasi-Orthogonality
Sewoong Lee, Adam Davies, Marc E. Canby +1
Sparse autoencoders (SAEs) are widely used in mechanistic interpretability research for large language models; however, the state-of-the-art method of using -sparse autoencoders…
cs.AI2024
Social Science Is Necessary for Operationalizing Socially Responsible Foundation Models
Adam Davies, Elisa Nguyen, Michael Simeone +2
With the rise of foundation models, there is growing concern about their potential social impacts. Social science has a long history of studying the social impacts of transformativ…