4 papers
JailMeter: An Evidence-Based Evaluation Framework for Jailbreak Attacks on Large Language Models
Qingjia Huang, Jingyu Zhang, Jianguo Wu +6
The assessment of jailbreak attacks against large language models currently suffers from inconsistent evaluation criteria and methods, leading to unreliable estimates of attack suc…
PRISM: Programmatic Reasoning with Image Sequence Manipulation for LVLM Jailbreaking
Quanchen Zou, Zonghao Ying, Moyang Chen +7
The increasing sophistication of large vision-language models (LVLMs) has been accompanied by advances in safety alignment mechanisms designed to prevent harmful content generation…
Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models
Yakai Li, Jiekang Hu, Weiduan Sang +7
Large Language Models face security threats from jailbreak attacks. Existing research has predominantly focused on prompt-level attacks while largely ignoring the underexplored att…
Latent Knowledge Scalpel: Precise and Massive Knowledge Editing for Large Language Models
Xin Liu, Qiyang Song, Shaowen Xu +6
Large Language Models (LLMs) often retain inaccurate or outdated information from pre-training, leading to incorrect predictions or biased outputs during inference. While existing…