9 citations · 9 across the 7 of their papers we have counts for
5 papers · 1 filter
Targeted Recovery of Weight-Space Mechanisms From Neural Networks
Antoine Vigouroux, Lee Sharkey
Parameter decomposition (PD) decomposes neural networks into interpretable computational components that faithfully reflect the original network's operations. However, scaling PD t…
Stochastic Parameter Decomposition
Lucius Bushnaq, Dan Braun, Lee Sharkey
A key step in reverse engineering neural networks is to decompose them into simpler parts that can be studied in relative isolation. Linear parameter decomposition -- a framework t…
Identifying Sparsely Active Circuits Through Local Loss Landscape Decomposition
Brianna Chrisman, Lucius Bushnaq, Lee Sharkey
Much of mechanistic interpretability has focused on understanding the activation spaces of large neural networks. However, activation space-based approaches reveal little about the…
Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition
Dan Braun, Lucius Bushnaq, Stefan Heimersheim +2
Mechanistic interpretability aims to understand the internal mechanisms learned by neural networks. Despite recent progress toward this goal, it remains unclear how best to decompo…
Open Problems in Mechanistic Interpretability
Lee Sharkey, Bilal Chughtai, Joshua Batson +26
Mechanistic interpretability aims to understand the computational mechanisms underlying neural networks' capabilities in order to accomplish concrete scientific and engineering goa…