2 papers
cs.AI2026
Evaluating whether AI models would sabotage AI safety research
Robert Kirk, Alexandra Souly, Kai Fronsdal +2
We evaluate the propensity of frontier models to sabotage or refuse to assist with safety research when deployed as AI research agents within a frontier AI company. We apply two co…
cs.AI2026
UK AISI Alignment Evaluation Case-Study
Alexandra Souly, Robert Kirk, Jacob Merizian +2
This technical report presents methods developed by the UK AI Security Institute for assessing whether advanced AI systems reliably follow intended goals. Specifically, we evaluate…