5 papers
Auditing Games for Sandbagging
Jordan Taylor, Sid Black, Dillon Bowen +10
Future AI systems could conceal their capabilities ('sandbagging') during evaluations, potentially misleading developers and auditors. We stress-tested sandbagging detection techni…
Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility
Brendan Murphy, Dillon Bowen, Shahrad Mohammadzadeh +4
AI systems are rapidly advancing in capability, and frontier model developers broadly acknowledge the need for safeguards against serious misuse. However, this paper demonstrates t…
Scaling Trends for Data Poisoning in LLMs
Dillon Bowen, Brendan Murphy, Will Cai +3
LLMs produce harmful and undesirable behavior when trained on datasets containing even a small fraction of poisoned data. We demonstrate that GPT models remain vulnerable to fine-t…
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models
Ann-Kathrin Dombrowski, Dillon Bowen, Adam Gleave +1
Open-weight large language models (LLMs) unlock huge benefits in innovation, personalization, privacy, and democratization. However, their core advantage - modifiability - opens th…
AI Companies Should Report Pre- and Post-Mitigation Safety Evaluations
Dillon Bowen, Ann-Kathrin Dombrowski, Adam Gleave +1
The rapid advancement of AI systems has raised widespread concerns about potential harms of frontier AI systems and the need for responsible evaluation and oversight. In this posit…