1 paper
Zhenhao Xu, Wenhan Chang, Yichuan Chen +3
Large Reasoning Models (LRMs) improve performance on complex tasks, but they also make safety control harder at deployment time. In black-box settings, defenders cannot modify mode…