1 paper
Zhenyu Wu, Siyuan Chen, Changchun Yang +8
Large Reasoning Models (LRMs) generate intermediate reasoning traces that may contain unsafe content, even when their final responses appear safe. Guardrail models are designed to…