3 papers
cs.AI2026
On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment
Yongjian Guo, Wanlun Ma, Lingyu Shen +2
Fine-tuning is the dominant paradigm for specializing large language models (LLMs), yet it exposes a critical vulnerability: malicious data providers can embed harmful behaviors in…
cs.CR2026
Beyond Her: Safety Dynamics in Role-play AI Companions
Zehang Deng, Zhaoyang Xie, Changzhou Han +8
The film 'Her' pictured a future of love between humans and AI. That future has quietly emerged in the form of Role-play AI Companions (RACs), where emotionally responsive interact…
cs.CR2020
Analysis of Trending Topics and Text-based Channels of Information Delivery in Cybersecurity
Tingmin Wu, Wanlun Ma, Sheng Wen +4
Computer users are generally faced with difficulties in making correct security decisions. While an increasingly fewer number of people are trying or willing to take formal securit…