6 citations · 6 across the 7 of their papers we have counts for
15 papers
Evaluating whether AI models would sabotage AI safety research
Robert Kirk, Alexandra Souly, Kai Fronsdal +2
We evaluate the propensity of frontier models to sabotage or refuse to assist with safety research when deployed as AI research agents within a frontier AI company. We apply two co…
UK AISI Alignment Evaluation Case-Study
Alexandra Souly, Robert Kirk, Jacob Merizian +2
This technical report presents methods developed by the UK AI Security Institute for assessing whether advanced AI systems reliably follow intended goals. Specifically, we evaluate…
How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition
Mateusz Dziemian, Maxwell Lin, Xiaohan Fu +28
LLM based agents are increasingly deployed in high stakes settings where they process external data sources such as emails, documents, and code repositories. This creates exposure…
Boundary Point Jailbreaking of Black-Box LLMs
Xander Davies, Giorgi Giglemiani, Edmund Lau +3
Frontier LLMs are safeguarded against attempts to extract harmful information via adversarial prompts known as "jailbreaks". Recently, defenders have developed classifier-based sys…
RedCodeAgent: Automatic Red-teaming Agent against Diverse Code Agents
Chengquan Guo, Chulin Xie, Yu Yang +6
Code agents have gained widespread adoption due to their strong code generation capabilities and integration with code interpreters, enabling dynamic execution, debugging, and inte…
Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
Alexandra Souly, Javier Rando, Ed Chapman +10
Poisoning attacks can compromise the safety of large language models (LLMs) by injecting malicious documents into their training data. Existing work has studied pretraining poisoni…