activity
20242026
collaborators

9 papers

cs.LG2026

Direction-Flipped Influence Audits Reveal Hidden Structure in Moral Choices of LLMs

Phil Blandfort, Tushar Karayil, Alex McKenzie +3

Moral benchmarks for LLMs typically score models on context-free prompts, implicitly treating the measured choice rate as stable. We test this assumption with a direction-flipped i…

cs.AI2026

The Impact of Off-Policy Training Data on Probe Generalisation

Nathalie Kirch, Samuel Dower, Adrians Skapars +3

Probing has emerged as a promising method for monitoring large language models (LLMs), enabling cheap inference-time detection of concerning behaviours. However, natural examples o…

cs.LG2026

Detecting High-Stakes Interactions with Activation Probes

Alex McKenzie, Urja Pawar, Phil Blandfort +4

Monitoring is an important aspect of safely deploying Large Language Models (LLMs). This paper examines activation probes for detecting ``high-stakes'' interactions -- where the te…

cs.LG2025

Fresh in memory: Training-order recency is linearly encoded in language model activations

Dmitrii Krasheninnikov, Richard E. Turner, David Krueger

We show that language models' activations linearly encode when information was learned during training. Our setup involves creating a model with a known training order by sequentia…

cs.LG2025

Understanding (Un)Reliability of Steering Vectors in Language Models

Joschka Braun, Carsten Eickhoff, David Krueger +2

Steering vectors are a lightweight method to control language model behavior by adding a learned bias to the activations at inference time. Although steering demonstrates promising…

cs.LG2025

Defining and Characterizing Reward Hacking

Joar Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov +1

We provide the first formal definition of reward hacking, a phenomenon where optimizing an imperfect proxy reward function leads to poor performance according to the true reward fu…