MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots
arXiv:2307.08715 · doi:10.14722/ndss.2024.24188
Abstract
Large Language Models (LLMs) have revolutionized Artificial Intelligence (AI) services due to their exceptional proficiency in understanding and generating human-like text. LLM chatbots, in particular, have seen widespread adoption, transforming human-machine interactions. However, these LLM chatbots are susceptible to "jailbreak" attacks, where malicious users manipulate prompts to elicit inappropriate or sensitive responses, contravening service policies. Despite existing attempts to mitigate such threats, our research reveals a substantial gap in our understanding of these vulnerabilities, largely due to the undisclosed defensive measures implemented by LLM service providers. In this paper, we present Jailbreaker, a comprehensive framework that offers an in-depth understanding of jailbreak attacks and countermeasures. Our work makes a dual contribution. First, we propose an innovative methodology inspired by time-based SQL injection techniques to reverse-engineer the defensive strategies of prominent LLM chatbots, such as ChatGPT, Bard, and Bing Chat. This time-sensitive approach uncovers intricate details about these services' defenses, facilitating a proof-of-concept attack that successfully bypasses their mechanisms. Second, we introduce an automatic generation method for jailbreak prompts. Leveraging a fine-tuned LLM, we validate the potential of automated jailbreak generation across various commercial LLM chatbots. Our method achieves a promising average success rate of 21.58%, significantly outperforming the effectiveness of existing techniques. We have responsibly disclosed our findings to the concerned service providers, underscoring the urgent need for more robust defenses. Jailbreaker thus marks a significant step towards understanding and mitigating jailbreak threats in the realm of LLM chatbots.
References in corpus (23)
- LLaMA: Open and Efficient Foundation Language Models
- A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- Large Language Models Are Human-Level Prompt Engineers
- Large Language Models for Software Engineering: A Systematic Literature Review
- Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study
- Prompt Injection attack against LLM-integrated Applications
- Ignore Previous Prompt: Attack Techniques For Language Models
- Fundamental Limitations of Alignment in Large Language Models
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models
- Prompting AI Art: An Investigation into the Creative Skill of Prompt Engineering
- RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment
- PentestGPT: An LLM-empowered Automatic Penetration Testing Tool
- Sources of Hallucination by Large Language Models on Inference Tasks
- Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks
- How Can ChatGPT Support Human Security Testers to Help Mitigate Supply Chain Attacks?
- Contrastive Learning Reduces Hallucination in Conversations
- Why So Toxic? Measuring and Triggering Toxic Behavior in Open-Domain Chatbots
- LLMs Can Understand Encrypted Prompt: Towards Privacy-Computing Friendly Transformers
- LMFlow: An Extensible Toolkit for Finetuning and Inference of Large Foundation Models
- Fine-mixing: Mitigating Backdoors in Fine-tuned Language Models
- NOTABLE: Transferable Backdoor Attacks Against Prompt-based NLP Models
- Automatic Prompt Optimization with "Gradient Descent" and Beam Search
Cited by in corpus (7)
- A Survey on Large Language Model (LLM) Security and Privacy: The Good, the Bad, and the Ugly
- AI Agents Under Threat: A Survey of Key Security Challenges and Future Pathways
- JailbreakLens: Visual Analysis of Jailbreak Attacks Against Large Language Models
- Improving LLM Outputs Against Jailbreak Attacks with Expert Model Integration
- GlitchProber: Advancing Effective Detection and Mitigation of Glitch Tokens in Large Language Models
- Helping Large Language Models Protect Themselves: An Enhanced Filtering and Summarization System
- CoCoTen: Detecting Adversarial Inputs to Large Language Models through Latent Space Features of Contextual Co-occurrence Tensors