collaborators

14 papers

cs.LG2026

Compressed Computation under Loss is likely Computation in Superposition

Francisco Ferreira da Silva, Stefan Heimersheim

Neural networks are thought to represent concepts as directions in their activation space, and superposition lets them encode more concepts than they have dimensions. It is natural…

cs.LG2026

Individual Parameters in Weight-Sparse Transformers Appear Interpretable

Arnau Marin-Llobet, Stefan Heimersheim

A central goal of mechanistic interpretability is to understand how neural networks work and what each individual component does. Dominant circuit-finding approaches focus on a spe…

cs.LG2026

The Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training Methodology

Andrzej Szablewski, Gabriel Konar-Steenberg, Raffaello Fornasiere +2

Model organisms (MOs) - language models trained to exhibit undesired or unnatural behaviours - are frequently used as testbeds for evaluating white-box interpretability techniques.…

cs.LG2026

Evidence for feature-specific error correction in LLMs

Francisco Ferreira da Silva, Stefan Heimersheim

Understanding the features of large language models (LLMs) is a central goal of interpretability. LLMs are commonly assumed to use superposition to represent more features than the…

cs.LG2026

Compressed Computation is (probably) not Computation in Superposition

Jai Bhagat, Sara Molas-Medina, Giorgi Giglemiani +1

We study whether the Compressed Computation (CC) toy model (Braun et al., 2025) is an instance of computation in superposition. The CC model appears to compute 100 ReLU functions w…

cs.LG2026

The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes

Mohammad Taufeeque, Stefan Heimersheim, Adam Gleave +1

Training against white-box deception detectors has been proposed as a way to make AI systems honest. However, such training risks models learning to obfuscate their deception to ev…