Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
EvoHarmBench: Breaking Content Moderation with Iterative Human-Like Evasion
Ruijie Jian, Benlei Cui, Ting Ma +8
Existing evaluations of harmful content detection rely predominantly on static benchmarks, which struggle to reflect the interactive adversarial ecosystem of real-world content pla…
cs.CL2026
Yuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety
Ting Ma, Xiufeng Huang, Benlei Cui +43
As large language models are increasingly deployed in real-world systems, safety failures can still lead to harmful outputs and dangerous misuse. We argue that the essence of safet…