1 paper
Chun Hei Yip, Rajashree Agrawal, Lawrence Chan +1
The goal of mechanistic interpretability is discovering simpler, low-rank algorithms implemented by models. While we can compress activations into features, compressing nonlinear f…