3 papers
cs.AI2026
SafeAtlas-VL: Beyond Binary Multimodal Safety with Large-Scale Data and Guard Models
Zongrui Wang, Xiangyang Zhu, Sicheng Wang +13
Multimodal safety moderation requires distinguishing risks arising from visual content, user intent, and assistant behavior. Existing safeguards, however, are typically trained for…
cs.CL2026
Emergent Misalignment Can Be Induced by Sycophancy and Reversed via Alignment Gating
Sicheng Wang, Xiangyang Zhu, Han Wang +6
Prior work has shown that fine-tuning large language models on malicious or incorrect outputs in narrow domains can induce broad misalignment and harmful behavior, a phenomenon kno…
cs.CL2026
Distinguishing Right from Wrong in Debates: Attribution Analysis of Chinese Harmful Memes
Weiming Wang, Junyu Lu, Han Wang +5
Research on harmful meme detection has garnered significant attention, resulting in the development of numerous datasets and methods. However, progress in detecting Chinese harmful…