Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Learning to Conceal Risk: Controllable Multi-turn Red Teaming for LLMs in the Financial Domain
Gang Cheng, Haibo Jin, Wenbin Zhang +2
Large Language Models (LLMs) are increasingly deployed in finance, where unsafe behavior can lead to serious regulatory risks. However, most red-teaming research focuses on overtly…
cs.CL2025
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation
Jun Zhuang, Haibo Jin, Ye Zhang +4
Intent detection, a core component of natural language understanding, has considerably evolved as a crucial mechanism in safeguarding large language models (LLMs). While prior work…