4 papers
Classifier Context Rot: Monitor Performance Degrades with Context Length
Sam Martin, Fabien Roger
Monitoring coding agents for dangerous behavior using language models requires classifying transcripts that often exceed 500K tokens, but prior agent monitoring benchmarks rarely c…
How Useful Is Cross-Domain Generalization for Training LLM Monitors?
Sam Martin, Fabien Roger
Using prompted language models as classifiers enables classification in domains with limited training data, but misses some of the robustness and performance benefits that fine-tun…
CTRL-ALT-DECEIT: Sabotage Evaluations for Automated AI R&D
Francis Rhys Ward, Teun van der Weij, Hanna Gábor +6
AI systems are increasingly able to autonomously conduct realistic software engineering tasks, and may soon be deployed to automate machine learning (ML) R&D itself. Frontier AI sy…
International Agreements on AI Safety: Review and Recommendations for a Conditional AI Safety Treaty
Rebecca Scholefield, Samuel Martin, Otto Barten
The malicious use or malfunction of advanced general-purpose AI (GPAI) poses risks that, according to leading experts, could lead to the 'marginalisation or extinction of humanity.…