Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
From Circuit Evidence to Mechanistic Theory: An Inductive Logic Approach
Nura Aljaafari, Danilo S. Carvalho, Andre Freitas
Mechanistic interpretability produces circuit-level causal analyses of neural network behaviour, but discovered circuits often remain isolated experimental artefacts: there is no s…
cs.LG2024
Transformer Normalisation Layers and the Independence of Semantic Subspaces
Stephen Menary, Samuel Kaski, Andre Freitas
Recent works have shown that transformers can solve contextual reasoning tasks by internally executing computational graphs called circuits. Circuits often use attention to logical…