2 papers
cs.CR2025
Adversarial Distilled Retrieval-Augmented Guarding Model for Online Malicious Intent Detection
Yihao Guo, Haocheng Bian, Liutong Zhou +16
With the deployment of Large Language Models (LLMs) in interactive applications, online malicious intent detection has become increasingly critical. However, existing approaches fa…
cs.CL2024
Playing Language Game with LLMs Leads to Jailbreaking
Yu Peng, Zewen Long, Fangming Dong +3
The advent of large language models (LLMs) has spurred the development of numerous jailbreak techniques aimed at circumventing their security defenses against malicious attacks. An…