14 papers
Compressed Computation under Loss is likely Computation in Superposition
Francisco Ferreira da Silva, Stefan Heimersheim
Neural networks are thought to represent concepts as directions in their activation space, and superposition lets them encode more concepts than they have dimensions. It is natural…
Individual Parameters in Weight-Sparse Transformers Appear Interpretable
Arnau Marin-Llobet, Stefan Heimersheim
A central goal of mechanistic interpretability is to understand how neural networks work and what each individual component does. Dominant circuit-finding approaches focus on a spe…
The Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training Methodology
Andrzej Szablewski, Gabriel Konar-Steenberg, Raffaello Fornasiere +2
Model organisms (MOs) - language models trained to exhibit undesired or unnatural behaviours - are frequently used as testbeds for evaluating white-box interpretability techniques.…
Evidence for feature-specific error correction in LLMs
Francisco Ferreira da Silva, Stefan Heimersheim
Understanding the features of large language models (LLMs) is a central goal of interpretability. LLMs are commonly assumed to use superposition to represent more features than the…
Compressed Computation is (probably) not Computation in Superposition
Jai Bhagat, Sara Molas-Medina, Giorgi Giglemiani +1
We study whether the Compressed Computation (CC) toy model (Braun et al., 2025) is an instance of computation in superposition. The CC model appears to compute 100 ReLU functions w…
The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes
Mohammad Taufeeque, Stefan Heimersheim, Adam Gleave +1
Training against white-box deception detectors has been proposed as a way to make AI systems honest. However, such training risks models learning to obfuscate their deception to ev…