1 paper
Yuhang Wang, Yanxu Zhu, Jiaming Zhang +2
Explicit safety policies can improve reasoning-model safety, but their effective coverage may lag behind evolving jailbreak strategies. We study whether a reasoning model can synth…