1 paper · 1 filter
Nura Aljaafari, Danilo S. Carvalho, Andre Freitas
Mechanistic interpretability produces circuit-level causal analyses of neural network behaviour, but discovered circuits often remain isolated experimental artefacts: there is no s…