1 paper
Angela van Sprang, Erman Acar, Willem Zuidema
Mechanistic interpretability focuses on reverse engineering the internal mechanisms learned by neural networks. We extend our focus and propose to mechanistically forward engineer…