19 citations · 27 across the 6 of their papers we have counts for
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
One Jailbreak, Many Tongues: Learning Language-Insensitive Intention Representations for Multilingual Jailbreak Detection
Shuyu Jiang, Kaiyu Xu, Xingshu Chen +5
Large language models (LLMs) are increasingly deployed in applications for global multilingual users, yet safety training remains concentrated in dominant languages and has not pro…
cs.CL2023★ 7 cited
Prompt Packer: Deceiving LLMs through Compositional Instruction with Hidden Attacks
Shuyu Jiang, Xingshu Chen, Rui Tang
Recently, Large language models (LLMs) with powerful general capabilities have been increasingly integrated into various Web applications, while undergoing alignment training to en…
cs.CL2023★ 1 cited
ReZG: Retrieval-Augmented Zero-Shot Counter Narrative Generation for Hate Speech
Shuyu Jiang, Wenyi Tang, Xingshu Chen +3
The proliferation of hate speech (HS) on social media poses a serious threat to societal security. Automatic counter narrative (CN) generation, as an active strategy for HS interve…