3 papers
cs.LG2026
The Ends Justify the Thoughts: RL-Induced Motivated Reasoning in LLM CoTs
Nikolaus Howe, Micah Carroll
Chain-of-Thought (CoT) monitoring has emerged as a compelling method for detecting harmful behaviors such as reward hacking for reasoning models, under the assumption that models'…
cs.LG2025
Scaling Trends in Language Model Robustness
Nikolaus Howe, Ian McKenzie, Oskar Hollinsworth +5
Increasing model size has unlocked a dazzling array of capabilities in modern language models. At the same time, even frontier models remain vulnerable to jailbreaks and prompt inj…
cs.LG2025
Defining and Characterizing Reward Hacking
Joar Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov +1
We provide the first formal definition of reward hacking, a phenomenon where optimizing an imperfect proxy reward function leads to poor performance according to the true reward fu…