Publications (9)
Verbalizable Representations Form a Global Workspace in Language Models
Wes Gurnee, Nicholas Sofroniew, Adam Pearce +13
Out of everything the human brain processes, only a small fraction is consciously accessible, in the sense of being available for verbal report, deliberate control, and flexible re…
Emergent misalignment as prompt sensitivity: A research note
Tim Wyse, Twm Stone, Anna Soligo +1
Betley et al. (2025) find that language models finetuned on insecure code become emergently misaligned (EM), giving misaligned responses in broad settings very different from those…
Convergent Linear Representations of Emergent Misalignment
Anna Soligo, Edward Turner, Senthooran Rajamanoharan +1
Fine-tuning large language models on narrow datasets can cause them to develop broadly misaligned behaviours: a phenomena known as emergent misalignment. However, the mechanisms un…
Model Organisms for Emergent Misalignment
Edward Turner, Anna Soligo, Mia Taylor +2
Recent work discovered Emergent Misalignment (EM): fine-tuning large language models on narrowly harmful datasets can lead them to become broadly misaligned. A survey of experts pr…
Gemma Needs Help: Investigating and Mitigating Emotional Instability in LLMs
Anna Soligo, Vladimir Mikulik, William Saunders
Large language models can generate responses that resemble emotional distress, and this raises concerns around model reliability and safety. We introduce a set of evaluations to in…
Interactions Between Crosscoder Features: A Compact Proofs Perspective
Dmitry Manning-Coe, Thomas Read, Anna Soligo +4
Dictionary learning methods like Sparse Autoencoders (SAEs) and crosscoders attempt to explain a model by decomposing its activations into independent features. Interactions betwee…