Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Bilinear MLPs enable weight-based mechanistic interpretability
Michael T. Pearce, Thomas Dooms, Alice Rigg +2
A mechanistic understanding of how MLPs do computation in deep neural networks remains elusive. Current interpretability work can extract features from hidden activations over an i…
cs.LG2024
Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations with MDL-SAEs
Kola Ayonrinde, Michael T. Pearce, Lee Sharkey
Sparse Autoencoders (SAEs) have emerged as a useful tool for interpreting the internal representations of neural networks. However, naively optimising SAEs for reconstruction loss…