Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
All Code, No Thought: Current Language Models Struggle to Reason in Ciphered Language
Shiyuan Guo, Henry Sleight, Fabien Roger
Detecting harmful AI actions is important as AI agents gain adoption. Chain-of-thought (CoT) monitoring is one method widely used to detect adversarial attacks and AI misalignment.…
cs.CL2025
Reasoning Models Don't Always Say What They Think
Yanda Chen, Joe Benton, Ansh Radhakrishnan +12
Chain-of-thought (CoT) offers a potential boon for AI safety as it allows monitoring a model's CoT to try to understand its intentions and reasoning processes. However, the effecti…