1 paper
Feitong Qiao, Liren Peng, Shiming Ren +7
Frontier language models that refuse harmful single-turn prompts often comply when the same intent is reached gradually over many turns, making multi-turn attacks one of the least…