2 papers
cs.CL2026
The Straight and Narrow: Do LLMs Possess an Internal Moral Path?
Luoming Hu, Jingjie Zeng, Liang Yang +1
Enhancing the moral alignment of Large Language Models (LLMs) is a critical challenge in AI safety. Current alignment techniques often act as superficial guardrails, leaving the in…
cs.CL2025
STATE ToxiCN: A Benchmark for Span-level Target-Aware Toxicity Extraction in Chinese Hate Speech Detection
Zewen Bai, Shengdi Yin, Junyu Lu +5
The proliferation of hate speech has caused significant harm to society. The intensity and directionality of hate are closely tied to the target and argument it is associated with.…