1 paper · 1 filter
Jacob Drori, Luke Marks, Bryce Woodworth +2
OpenAI (2025) showed that training against a chain of thought (CoT) monitor can cause obfuscated CoTs, which contain bad behavior the monitor cannot detect. They proposed to keep C…