1 paper · 1 filter
Artyom Mazur, Nina Konovalova, Aibek Alanov
Mechanistic interpretability seeks to explain neural network behavior by decomposing model computations into interpretable features and circuits. While transcoder-based circuit tra…