A Security Risk Taxonomy for Prompt-Based Interaction With Large Language Models
arXiv:2311.11415 · doi:10.1109/ACCESS.2024.3450388
Abstract
As large language models (LLMs) permeate more and more applications, an assessment of their associated security risks becomes increasingly necessary. The potential for exploitation by malicious actors, ranging from disinformation to data breaches and reputation damage, is substantial. This paper addresses a gap in current research by specifically focusing on security risks posed by LLMs within the prompt-based interaction scheme, which extends beyond the widely covered ethical and societal implications. Our work proposes a taxonomy of security risks along the user-model communication pipeline and categorizes the attacks by target and attack type alongside the commonly used confidentiality, integrity, and availability (CIA) triad. The taxonomy is reinforced with specific attack examples to showcase the real-world impact of these risks. Through this taxonomy, we aim to inform the development of robust and secure LLM applications, enhancing their safety and trustworthiness.
References in corpus (10)
- Evaluating Large Language Models Trained on Code
- BOLD: Dataset and Metrics for Measuring Biases in Open-Ended Language Generation
- Understanding the Capabilities, Limitations, and Societal Impact of Large Language Models
- Beyond the Safeguards: Exploring the Security Risks of ChatGPT
- From Text to MITRE Techniques: Exploring the Malicious Use of Large Language Models for Generating Cyber Attack Payloads
- BadGPT: Exploring Security Vulnerabilities of ChatGPT via Backdoor Attacks to InstructGPT
- Safety Assessment of Chinese Large Language Models
- LLMs Killed the Script Kiddie: How Agents Supported by Large Language Models Change the Landscape of Network Threat Testing
- Assessing Language Model Deployment with Risk Cards
- Trainwreck: A damaging adversarial attack on image classifiers