1 paper · 1 filter
JungMin Yun, Junehyoung Kwon, Hayeong Ryu +3
Large Reasoning Models (LRMs) pose a dual-surface safety challenge: both intermediate reasoning traces and final answers can contain harmful content. Existing alignment methods oft…