3 papers
cs.LG2026
Detecting High-Stakes Interactions with Activation Probes
Alex McKenzie, Urja Pawar, Phil Blandfort +4
Monitoring is an important aspect of safely deploying Large Language Models (LLMs). This paper examines activation probes for detecting ``high-stakes'' interactions -- where the te…
cs.LG2025
Fresh in memory: Training-order recency is linearly encoded in language model activations
Dmitrii Krasheninnikov, Richard E. Turner, David Krueger
We show that language models' activations linearly encode when information was learned during training. Our setup involves creating a model with a known training order by sequentia…
cs.LG2025
Understanding (Un)Reliability of Steering Vectors in Language Models
Joschka Braun, Carsten Eickhoff, David Krueger +2
Steering vectors are a lightweight method to control language model behavior by adding a learned bias to the activations at inference time. Although steering demonstrates promising…