Showing 2024Show all
3 papers · 1 filter
cs.CR2024
Unleashing the Unseen: Harnessing Benign Datasets for Jailbreaking Large Language Models
Wei Zhao, Zhe Li, Yige Li +1
Despite significant ongoing efforts in safety alignment, large language models (LLMs) such as GPT-4 and LLaMA 3 remain vulnerable to jailbreak attacks that can induce harmful behav…
cs.CL2024
Do Influence Functions Work on Large Language Models?
Zhe Li, Wei Zhao, Yige Li +1
Influence functions are important for quantifying the impact of individual training data points on a model's predictions. Although extensive research has been conducted on influenc…
cs.AI2024
Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing
Wei Zhao, Zhe Li, Yige Li +2
Large language models (LLMs) are increasingly being adopted in a wide range of real-world applications. Despite their impressive performance, recent studies have shown that LLMs ar…