3 papers
cs.AI2026
Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring
Jennifer Za, Julija Bainiaksina, Nikita Ostrovsky +2
Chain-of-thought (CoT) monitoring is a promising safety mechanism for AI agents, based on the premise that visible reasoning traces can surface misaligned or deceptive behavior. Wh…
cs.LG2025
World Model Robustness via Surprise Recognition
Geigh Zollicoffer, Tanush Chopra, Mingkuan Yan +3
AI systems deployed in the real world must contend with distractions and out-of-distribution (OOD) noise that can destabilize their policies and lead to unsafe behavior. While robu…
cs.CL2024
View From Above: A Framework for Evaluating Distribution Shifts in Model Behavior
Tanush Chopra, Michael Li, Jacob Haimes
When large language models (LLMs) are asked to perform certain tasks, how can we be sure that their learned representations align with reality? We propose a domain-agnostic framewo…