Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds
Axel Ahlqvist, Richard Guan, Juan-Pablo Rivera +6
A core obstacle to alignment evaluation is evaluation awareness: capable models can tell when they are being tested rather than deployed, weakening the conclusions a safety evaluat…
cs.AI2024
Escalation Risks from Language Models in Military and Diplomatic Decision-Making
Juan-Pablo Rivera, Gabriel Mukobi, Anka Reuel +3
Governments are increasingly considering integrating autonomous AI agents in high-stakes military and foreign-policy decision-making, especially with the emergence of advanced gene…