1 paper · 1 filter
Mikita Balesni, Tomek Korbak, Owain Evans
Large language models can use chain-of-thought (CoT) to externalize reasoning, potentially enabling oversight of capable LLM agents. Prior work has shown that models struggle at tw…