Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation
Jun Zhuang, Haibo Jin, Ye Zhang +4
Intent detection, a core component of natural language understanding, has considerably evolved as a crucial mechanism in safeguarding large language models (LLMs). While prior work…
cs.CL2025
ChallengeMe: An Adversarial Learning-enabled Text Summarization Framework
Xiaoyu Deng, Ye Zhang, Tianmin Guo +3
The astonishing performance of large language models (LLMs) and their remarkable achievements in production and daily life have led to their widespread application in collaborative…