1 paper
Jiaming Ji, Wenqi Chen, Kaile Wang +8
Modern large language models rely on chain-of-thought (CoT) reasoning to achieve impressive performance, yet the same mechanism can amplify deceptive alignment, situations in which…