3 papers
cs.LG2025
How Reliable are Causal Probing Interventions?
Marc Canby, Adam Davies, Chirag Rastogi +1
Causal probing aims to analyze foundation models by examining how intervening on their representation of various latent properties impacts their outputs. Recent works have cast dou…
cs.LG2025
Evaluating and Designing Sparse Autoencoders by Approximating Quasi-Orthogonality
Sewoong Lee, Adam Davies, Marc E. Canby +1
Sparse autoencoders (SAEs) are widely used in mechanistic interpretability research for large language models; however, the state-of-the-art method of using -sparse autoencoders…
cs.AI2025
Social Science Is Necessary for Operationalizing Socially Responsible Foundation Models
Adam Davies, Elisa Nguyen, Michael Simeone +2
With the rise of foundation models, there is growing concern about their potential social impacts. Social science has a long history of studying the social impacts of transformativ…