Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Propensity Inference: Environmental Contributors to LLM Behaviour
Olli Järviniemi, Oliver Makins, Jacob Merizian +2
Motivated by loss of control risks from misaligned AI systems, we develop and apply methods for measuring language models' propensity for unsanctioned behaviour. We contribute thre…
cs.AI2025
How Do Large Language Monkeys Get Their Power (Laws)?
Rylan Schaeffer, Joshua Kazdan, John Hughes +7
Recent research across mathematical problem solving, proof assistant programming and multimodal jailbreaking documents a striking finding: when (multimodal) language model tackle a…