Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Emergent Misalignment Can Be Induced by Sycophancy and Reversed via Alignment Gating
Sicheng Wang, Xiangyang Zhu, Han Wang +6
Prior work has shown that fine-tuning large language models on malicious or incorrect outputs in narrow domains can induce broad misalignment and harmful behavior, a phenomenon kno…
cs.CL2026
Distinguishing Right from Wrong in Debates: Attribution Analysis of Chinese Harmful Memes
Weiming Wang, Junyu Lu, Han Wang +5
Research on harmful meme detection has garnered significant attention, resulting in the development of numerous datasets and methods. However, progress in detecting Chinese harmful…