4 papers
HumorReject: Decoupling LLM Safety from Refusal Prefix via A Little Humor
Zihui Wu, Haichang Gao, Jiacheng Luo +1
Large Language Models (LLMs) commonly rely on explicit refusal prefixes for safety, making them vulnerable to prefix injection attacks. We introduce HumorReject, a novel data-drive…
GlitchMiner: Mining Glitch Tokens in Large Language Models via Gradient-based Discrete Optimization
Zihui Wu, Haichang Gao, Ping Wang +3
Glitch tokens, inputs that trigger unpredictable or anomalous behavior in Large Language Models (LLMs), pose significant challenges to model reliability and safety. Existing detect…
Whispers of Data: Unveiling Label Distributions in Federated Learning Through Virtual Client Simulation
Zhixuan Ma, Haichang Gao, Junxiang Huang +1
Federated Learning enables collaborative training of a global model across multiple geographically dispersed clients without the need for data sharing. However, it is susceptible t…
The Dark Side of Function Calling: Pathways to Jailbreaking Large Language Models
Zihui Wu, Haichang Gao, Jianping He +1
Large language models (LLMs) have demonstrated remarkable capabilities, but their power comes with significant security considerations. While extensive research has been conducted…