Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
How Reliable are Causal Probing Interventions?
Marc Canby, Adam Davies, Chirag Rastogi +1
Causal probing aims to analyze foundation models by examining how intervening on their representation of various latent properties impacts their outputs. Recent works have cast dou…
cs.LG2025
Evaluating and Designing Sparse Autoencoders by Approximating Quasi-Orthogonality
Sewoong Lee, Adam Davies, Marc E. Canby +1
Sparse autoencoders (SAEs) are widely used in mechanistic interpretability research for large language models; however, the state-of-the-art method of using -sparse autoencoders…