3 papers
cs.AI2026
Propensity Inference: Environmental Contributors to LLM Behaviour
Olli Järviniemi, Oliver Makins, Jacob Merizian +2
Motivated by loss of control risks from misaligned AI systems, we develop and apply methods for measuring language models' propensity for unsanctioned behaviour. We contribute thre…
cs.LG2025
Analyzing the Generalization and Reliability of Steering Vectors
Daniel Tan, David Chanin, Aengus Lynch +4
Steering vectors (SVs) have been proposed as an effective approach to adjust language model behaviour at inference time by intervening on intermediate model activations. They have…
cs.AI2025
How Do Large Language Monkeys Get Their Power (Laws)?
Rylan Schaeffer, Joshua Kazdan, John Hughes +7
Recent research across mathematical problem solving, proof assistant programming and multimodal jailbreaking documents a striking finding: when (multimodal) language model tackle a…