18 citations · 21 across the 2 of their papers we have counts for
3 papers · 1 filter
Persona Features Control Emergent Misalignment
Miles Wang, Tom Dupré la Tour, Olivia Watkins +8
Understanding how language models generalize behaviors from their training to a broader deployment distribution is an important problem in AI safety. Betley et al. discovered that…
Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control
Aleksandar Makelov, George Lange, Neel Nanda
Disentangling model activations into meaningful features is a central problem in interpretability. However, the absence of ground-truth for these features in realistic scenarios ma…
Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation Patching
Aleksandar Makelov, Georg Lange, Neel Nanda
Mechanistic interpretability aims to understand model behaviors in terms of specific, interpretable features, often hypothesized to manifest as low-dimensional subspaces of activat…