1 paper
Nils Palumbo, Ravi Mangal, Zifan Wang +3
Mechanistic interpretability aims to reverse engineer the computation performed by a neural network in terms of its internal components. Although there is a growing body of researc…