2 papers
cs.AI2026
Propensity Inference: Environmental Contributors to LLM Behaviour
Olli Järviniemi, Oliver Makins, Jacob Merizian +2
Motivated by loss of control risks from misaligned AI systems, we develop and apply methods for measuring language models' propensity for unsanctioned behaviour. We contribute thre…
cs.CR2025
RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents
Sid Black, Asa Cooper Stickland, Jake Pencharz +7
Uncontrollable autonomous replication of language model agents poses a critical safety risk. To better understand this risk, we introduce RepliBench, a suite of evaluations designe…