1 paper
Shiyuan Guo, Henry Sleight, Fabien Roger
Detecting harmful AI actions is important as AI agents gain adoption. Chain-of-thought (CoT) monitoring is one method widely used to detect adversarial attacks and AI misalignment.…