1 paper
Casper L. Christensen, Logan Riggs
Recent work in mechanistic interpretability has shown that decomposing models in parameter space may yield clean handles for analysis and intervention. Previous methods have demons…