3 papers
cs.LG2026
Residual Stream Analysis of Overfitting And Structural Disruptions
Quan Liu, Han Zhou, Wenquan Wu +2
Ensuring that large language models (LLMs) remain both helpful and harmless poses a significant challenge: fine-tuning on repetitive safety datasets, where unsafe prompts are paire…
cs.CL2024
Alignment-Enhanced Decoding:Defending via Token-Level Adaptive Refining of Probability Distributions
Quan Liu, Zhenhong Zhou, Longzhu He +3
Large language models are susceptible to jailbreak attacks, which can result in the generation of harmful content. While prior defenses mitigate these risks by perturbing or inspec…
cs.CL2024
Speak Out of Turn: Safety Vulnerability of Large Language Models in Multi-turn Dialogue
Zhenhong Zhou, Jiuyang Xiang, Haopeng Chen +3
Large Language Models (LLMs) have been demonstrated to generate illegal or unethical responses, particularly when subjected to "jailbreak." Research on jailbreak has highlighted th…