10 citations · 10 across the 1 of their papers we have counts for
9 papers
Mitigating Goal Misgeneralization via Minimax Regret
Karim Abdel Sadek, Matthew Farrugia-Roberts, Usman Anwar +4
Safe generalization in reinforcement learning requires not only that a learned policy acts capably in new situations, but also that it uses its capabilities towards the pursuit of…
REAL: Benchmarking Autonomous Agents on Deterministic Simulations of Real Websites
Divyansh Garg, Shaun VanWeelden, Diego Caples +15
We introduce REAL, a benchmark and framework for multi-turn agent evaluations on deterministic simulations of real-world websites. REAL comprises high-fidelity, deterministic repli…
Multi-Agent Risks from Advanced AI
Lewis Hammond, Alan Chan, Jesse Clifton +41
The rapid development of advanced AI agents and the imminent deployment of many instances of these agents will give rise to multi-agent systems of unprecedented complexity. These s…
Fundamental Limitations in Pointwise Defences of LLM Finetuning APIs
Xander Davies, Eric Winsor, Alexandra Souly +4
LLM developers have imposed technical interventions to prevent fine-tuning misuse attacks, attacks where adversaries evade safeguards by fine-tuning the model using a public API. P…
Humanity's Last Exam
Long Phan, Alice Gatti, Ziwen Han +1144
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…
Delta-Influence: Unlearning Poisons via Influence Functions
Wenjie Li, Jiawei Li, Pengcheng Zeng +3
Addressing data integrity challenges, such as unlearning the effects of data poisoning after model training, is necessary for the reliable deployment of machine learning models. St…