2 papers
cs.AI2026
Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring
Jennifer Za, Julija Bainiaksina, Nikita Ostrovsky +2
Chain-of-thought (CoT) monitoring is a promising safety mechanism for AI agents, based on the premise that visible reasoning traces can surface misaligned or deceptive behavior. Wh…
cs.LG2025
World Model Robustness via Surprise Recognition
Geigh Zollicoffer, Tanush Chopra, Mingkuan Yan +3
AI systems deployed in the real world must contend with distractions and out-of-distribution (OOD) noise that can destabilize their policies and lead to unsafe behavior. While robu…