3 papers
cs.CR2026
Beyond Her: Safety Dynamics in Role-play AI Companions
Zehang Deng, Zhaoyang Xie, Changzhou Han +8
The film 'Her' pictured a future of love between humans and AI. That future has quietly emerged in the form of Role-play AI Companions (RACs), where emotionally responsive interact…
cs.LG2026
Sharpness-Aware Minimization in Logit Space Efficiently Enhances Direct Preference Optimization
Haocheng Luo, Zehang Deng, Thanh-Toan Do +3
Direct Preference Optimization (DPO) has emerged as a popular algorithm for aligning pretrained large language models with human preferences, owing to its simplicity and training s…
cs.CV2026
Decoupling Defense Strategies for Robust Image Watermarking
Jiahui Chen, Zehang Deng, Zeyu Zhang +3
Deep learning-based image watermarking, while robust against conventional distortions, remains vulnerable to advanced adversarial and regeneration attacks. Conventional countermeas…