1 paper
Brianna Chrisman, Lucius Bushnaq, Lee Sharkey
Much of mechanistic interpretability has focused on understanding the activation spaces of large neural networks. However, activation space-based approaches reveal little about the…