papers

Publications (9)

cs.CL2026

Verbalizable Representations Form a Global Workspace in Language Models

Wes Gurnee, Nicholas Sofroniew, Adam Pearce +13

Out of everything the human brain processes, only a small fraction is consciously accessible, in the sense of being available for verbal report, deliberate control, and flexible re…

cs.CR2025

Emergent misalignment as prompt sensitivity: A research note

Tim Wyse, Twm Stone, Anna Soligo +1

Betley et al. (2025) find that language models finetuned on insecure code become emergently misaligned (EM), giving misaligned responses in broad settings very different from those…

cs.LG2025

Convergent Linear Representations of Emergent Misalignment

Anna Soligo, Edward Turner, Senthooran Rajamanoharan +1

Fine-tuning large language models on narrow datasets can cause them to develop broadly misaligned behaviours: a phenomena known as emergent misalignment. However, the mechanisms un…

cs.LG2025

Model Organisms for Emergent Misalignment

Edward Turner, Anna Soligo, Mia Taylor +2

Recent work discovered Emergent Misalignment (EM): fine-tuning large language models on narrowly harmful datasets can lead them to become broadly misaligned. A survey of experts pr…

cs.CL2026

Gemma Needs Help: Investigating and Mitigating Emotional Instability in LLMs

Anna Soligo, Vladimir Mikulik, William Saunders

Large language models can generate responses that resemble emotional distress, and this raises concerns around model reliability and safety. We introduce a set of evaluations to in…

cs.LG2026

Interactions Between Crosscoder Features: A Compact Proofs Perspective

Dmitry Manning-Coe, Thomas Read, Anna Soligo +4

Dictionary learning methods like Sparse Autoencoders (SAEs) and crosscoders attempt to explain a model by decomposing its activations into independent features. Interactions betwee…