23 citations · 23 across the 4 of their papers we have counts for
4 papers
The Ends Justify the Thoughts: RL-Induced Motivated Reasoning in LLM CoTs
Nikolaus Howe, Micah Carroll
Chain-of-Thought (CoT) monitoring has emerged as a compelling method for detecting harmful behaviors such as reward hacking for reasoning models, under the assumption that models'…
Scaling Trends in Language Model Robustness
Nikolaus Howe, Ian McKenzie, Oskar Hollinsworth +5
Increasing model size has unlocked a dazzling array of capabilities in modern language models. At the same time, even frontier models remain vulnerable to jailbreaks and prompt inj…
Defining and Characterizing Reward Hacking
Joar Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov +1
We provide the first formal definition of reward hacking, a phenomenon where optimizing an imperfect proxy reward function leads to poor performance according to the true reward fu…
Myriad: a real-world testbed to bridge trajectory optimization and deep learning
Nikolaus H. R. Howe, Simon Dufort-Labbé, Nitarshan Rajkumar +1
We present Myriad, a testbed written in JAX for learning and planning in real-world continuous environments. The primary contributions of Myriad are threefold. First, Myriad provid…