3 citations · 3 across the 1 of their papers we have counts for
2 papers
cs.LG2024★ 3 cited
Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control
Aleksandar Makelov, George Lange, Neel Nanda
Disentangling model activations into meaningful features is a central problem in interpretability. However, the absence of ground-truth for these features in realistic scenarios ma…
cs.LG2023
Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation Patching
Aleksandar Makelov, Georg Lange, Neel Nanda
Mechanistic interpretability aims to understand model behaviors in terms of specific, interpretable features, often hypothesized to manifest as low-dimensional subspaces of activat…