JailbreakLens: Visual Analysis of Jailbreak Attacks Against Large Language Models
arXiv:2404.08793 · doi:10.1109/TVCG.2025.3575694
Abstract
The proliferation of large language models (LLMs) has underscored concerns regarding their security vulnerabilities, notably against jailbreak attacks, where adversaries design jailbreak prompts to circumvent safety mechanisms for potential misuse. Addressing these concerns necessitates a comprehensive analysis of jailbreak prompts to evaluate LLMs' defensive capabilities and identify potential weaknesses. However, the complexity of evaluating jailbreak performance and understanding prompt characteristics makes this analysis laborious. We collaborate with domain experts to characterize problems and propose an LLM-assisted framework to streamline the analysis process. It provides automatic jailbreak assessment to facilitate performance evaluation and support analysis of components and keywords in prompts. Based on the framework, we design JailbreakLens, a visual analysis system that enables users to explore the jailbreak performance against the target model, conduct multi-level analysis of prompt characteristics, and refine prompt instances to verify findings. Through a case study, technical evaluations, and expert interviews, we demonstrate our system's effectiveness in helping users evaluate model security and identify model weaknesses.
References in corpus (12)
- The What-If Tool: Interactive Probing of Machine Learning Models
- explAIner: A Visual Analytics Framework for Interactive and Explainable Machine Learning
- Sensecape: Enabling Multilevel Exploration and Sensemaking with Large Language Models
- Graphologue: Exploring Large Language Model Responses with Interactive Diagrams
- PromptMagician: Interactive Prompt Engineering for Text-to-Image Creation
- "What It Wants Me To Say": Bridging the Abstraction Gap Between End-User Programmers and Code-Generating Large Language Models
- MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots
- Explaining Vulnerabilities to Adversarial Machine Learning through Visual Analytics
- Spellburst: A Node-based Interface for Exploratory Creative Coding with Natural Language Prompts
- XNLI: Explaining and Diagnosing NLI-based Visual Data Analysis
- Storyfier: Exploring Vocabulary Learning Support with Text Generation Models
- CommonsenseVIS: Visualizing and Understanding Commonsense Reasoning Capabilities of Natural Language Models