Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Probing Social Identity Bias in Chinese LLMs with Gendered Pronouns and Social Groups
Geng Liu, Feng Li, Junjie Mu +2
Large language models (LLMs) are increasingly deployed in user-facing applications, raising concerns that they may reflect and amplify social biases. We investigate social identity…
cs.CL2026
Mask-GCG: Are All Tokens in Adversarial Suffixes Necessary for Jailbreak Attacks?
Junjie Mu, Zonghao Ying, Zhekui Fan +6
Jailbreak attacks on Large Language Models (LLMs) have demonstrated various successful methods whereby attackers manipulate models into generating harmful responses that they are d…