4 papers · 1 filter
Self-ReSET: Learning to Self-Recover from Unsafe Reasoning Trajectories
Dongcheng Zhang, Yi Zhang, Yuxin Chen +3
Large Reasoning Models possess remarkable capabilities for self-correction in general domain; however, they frequently struggle to recover from unsafe reasoning trajectories under…
Internalizing Safety Understanding in Large Reasoning Models via Verification
Yi Zhang, Yuxin Chen, Leheng Sheng +4
While explicit Chain-of-Thought (CoT) empowers large reasoning models (LRMs), it enables the generation of riskier final answers. Current alignment paradigms primarily rely on exte…
On Reasoning Strength Planning in Large Reasoning Models
Leheng Sheng, An Zhang, Zijian Wu +5
Recent studies empirically reveal that large reasoning models (LRMs) can automatically allocate more reasoning strengths (i.e., the number of reasoning tokens) for harder problems,…
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning
Yi Zhang, An Zhang, XiuYu Zhang +4
Large language models (LLMs), despite possessing latent safety understanding from their vast pretraining data, remain vulnerable to generating harmful content and exhibit issues su…