3 papers
cs.AI2026
Propensity Inference: Environmental Contributors to LLM Behaviour
Olli Järviniemi, Oliver Makins, Jacob Merizian +2
Motivated by loss of control risks from misaligned AI systems, we develop and apply methods for measuring language models' propensity for unsanctioned behaviour. We contribute thre…
cs.AI2025
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Elliot Glazer, Ege Erdil, Tamay Besiroglu +21
We introduce FrontierMath, a benchmark of hundreds of original, exceptionally challenging mathematics problems crafted and vetted by expert mathematicians. The questions cover most…
cs.CL2025
Subversion via Focal Points: Investigating Collusion in LLM Monitoring
Olli Järviniemi
We evaluate language models' ability to subvert monitoring protocols via collusion. More specifically, we have two instances of a model design prompts for a policy (P) and a monito…