9 papers
Direction-Flipped Influence Audits Reveal Hidden Structure in Moral Choices of LLMs
Phil Blandfort, Tushar Karayil, Alex McKenzie +3
Moral benchmarks for LLMs typically score models on context-free prompts, implicitly treating the measured choice rate as stable. We test this assumption with a direction-flipped i…
The Impact of Off-Policy Training Data on Probe Generalisation
Nathalie Kirch, Samuel Dower, Adrians Skapars +3
Probing has emerged as a promising method for monitoring large language models (LLMs), enabling cheap inference-time detection of concerning behaviours. However, natural examples o…
Detecting High-Stakes Interactions with Activation Probes
Alex McKenzie, Urja Pawar, Phil Blandfort +4
Monitoring is an important aspect of safely deploying Large Language Models (LLMs). This paper examines activation probes for detecting ``high-stakes'' interactions -- where the te…
Fresh in memory: Training-order recency is linearly encoded in language model activations
Dmitrii Krasheninnikov, Richard E. Turner, David Krueger
We show that language models' activations linearly encode when information was learned during training. Our setup involves creating a model with a known training order by sequentia…
Understanding (Un)Reliability of Steering Vectors in Language Models
Joschka Braun, Carsten Eickhoff, David Krueger +2
Steering vectors are a lightweight method to control language model behavior by adding a learned bias to the activations at inference time. Although steering demonstrates promising…
Defining and Characterizing Reward Hacking
Joar Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov +1
We provide the first formal definition of reward hacking, a phenomenon where optimizing an imperfect proxy reward function leads to poor performance according to the true reward fu…