1 citations · 1 across the 3 of their papers we have counts for
5 papers · 1 filter
Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition
Dan Braun, Lucius Bushnaq, Stefan Heimersheim +2
Mechanistic interpretability aims to understand the internal mechanisms learned by neural networks. Despite recent progress toward this goal, it remains unclear how best to decompo…
Mathematical Models of Computation in Superposition
Kaarel Hänni, Jake Mendel, Dmitry Vaintrob +1
Superposition -- when a neural network represents more ``features'' than it has dimensions -- seems to pose a serious challenge to mechanistically interpreting current AI systems.…
The Local Interaction Basis: Identifying Computationally-Relevant and Sparsely Interacting Features in Neural Networks
Lucius Bushnaq, Stefan Heimersheim, Nicholas Goldowsky-Dill +7
Mechanistic interpretability aims to understand the behavior of neural networks by reverse-engineering their internal computations. However, current methods struggle to find clear…
Using Degeneracy in the Loss Landscape for Mechanistic Interpretability
Lucius Bushnaq, Jake Mendel, Stefan Heimersheim +5
Mechanistic Interpretability aims to reverse engineer the algorithms implemented by neural networks by studying their weights and activations. An obstacle to reverse engineering ne…
Dynamical versus Bayesian Phase Transitions in a Toy Model of Superposition
Zhongtian Chen, Edmund Lau, Jake Mendel +2
We investigate phase transitions in a Toy Model of Superposition (TMS) using Singular Learning Theory (SLT). We derive a closed formula for the theoretical loss and, in the case of…