Showing 2024Show all
2 papers · 1 filter
cs.LG2024
Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations with MDL-SAEs
Kola Ayonrinde, Michael T. Pearce, Lee Sharkey
Sparse Autoencoders (SAEs) have emerged as a useful tool for interpreting the internal representations of neural networks. However, naively optimising SAEs for reconstruction loss…
cs.LG2024
Weight-based Decomposition: A Case for Bilinear MLPs
Michael T. Pearce, Thomas Dooms, Alice Rigg
Gated Linear Units (GLUs) have become a common building block in modern foundation models. Bilinear layers drop the non-linearity in the "gate" but still have comparable performanc…