Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Adaptive and Explicit safe: Triggering Latent Safety Awareness in Large Reasoning Models
Ke Miao, Jiaxin Li, Hongliang Chen +2
While Large Reasoning Models (LRMs) excel at complex tasks, they remain highly vulnerable to sophisticated jailbreaks and direct harmful queries. To address this vulnerability, pri…
cs.AI2025
Towards Evaluation for Real-World LLM Unlearning
Ke Miao, Yuke Hu, Xiaochen Li +4
This paper analyzes the limitations of existing unlearning evaluation metrics in terms of practicality, exactness, and robustness in real-world LLM unlearning scenarios. To overcom…