Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Self-ReSET: Learning to Self-Recover from Unsafe Reasoning Trajectories
Dongcheng Zhang, Yi Zhang, Yuxin Chen +3
Large Reasoning Models possess remarkable capabilities for self-correction in general domain; however, they frequently struggle to recover from unsafe reasoning trajectories under…
cs.AI2026
Internalizing Safety Understanding in Large Reasoning Models via Verification
Yi Zhang, Yuxin Chen, Leheng Sheng +4
While explicit Chain-of-Thought (CoT) empowers large reasoning models (LRMs), it enables the generation of riskier final answers. Current alignment paradigms primarily rely on exte…
cs.AI2024
ALI-Agent: Assessing LLMs' Alignment with Human Values via Agent-based Evaluation
Jingnan Zheng, Han Wang, An Zhang +3
Large Language Models (LLMs) can elicit unintended and even harmful content when misaligned with human values, posing severe risks to users and society. To mitigate these risks, cu…