Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Classifier Context Rot: Monitor Performance Degrades with Context Length
Sam Martin, Fabien Roger
Monitoring coding agents for dangerous behavior using language models requires classifying transcripts that often exceed 500K tokens, but prior agent monitoring benchmarks rarely c…
cs.AI2026
How Useful Is Cross-Domain Generalization for Training LLM Monitors?
Sam Martin, Fabien Roger
Using prompted language models as classifiers enables classification in domains with limited training data, but misses some of the robustness and performance benefits that fine-tun…
cs.AI2025
CTRL-ALT-DECEIT: Sabotage Evaluations for Automated AI R&D
Francis Rhys Ward, Teun van der Weij, Hanna Gábor +6
AI systems are increasingly able to autonomously conduct realistic software engineering tasks, and may soon be deployed to automate machine learning (ML) R&D itself. Frontier AI sy…