Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Evaluating and Designing Sparse Autoencoders by Approximating Quasi-Orthogonality
Sewoong Lee, Adam Davies, Marc E. Canby +1
Sparse autoencoders (SAEs) are widely used in mechanistic interpretability research for large language models; however, the state-of-the-art method of using -sparse autoencoders…
cs.LG2024
How Reliable are Causal Probing Interventions?
Marc Canby, Adam Davies, Chirag Rastogi +1
Causal probing aims to analyze foundation models by examining how intervening on their representation of various latent properties impacts their outputs. Recent works have cast dou…