A Survey on Large Language Model (LLM) Security and Privacy: The Good, the Bad, and the Ugly
arXiv:2312.02003 · doi:10.1016/j.hcc.2024.100211
Abstract
Large Language Models (LLMs), such as ChatGPT and Bard, have revolutionized natural language understanding and generation. They possess deep language comprehension, human-like text generation capabilities, contextual awareness, and robust problem-solving skills, making them invaluable in various domains (e.g., search engines, customer support, translation). In the meantime, LLMs have also gained traction in the security community, revealing security vulnerabilities and showcasing their potential in security-related tasks. This paper explores the intersection of LLMs with security and privacy. Specifically, we investigate how LLMs positively impact security and privacy, potential risks and threats associated with their use, and inherent vulnerabilities within LLMs. Through a comprehensive literature review, the paper categorizes the papers into "The Good" (beneficial LLM applications), "The Bad" (offensive applications), and "The Ugly" (vulnerabilities of LLMs and their defenses). We have some interesting findings. For example, LLMs have proven to enhance code security (code vulnerability detection) and data privacy (data confidentiality protection), outperforming traditional methods. However, they can also be harnessed for various attacks (particularly user-level attacks) due to their human-like reasoning abilities. We have identified areas that require further research efforts. For example, Research on model and parameter extraction attacks is limited and often theoretical, hindered by LLM parameter scale and confidentiality. Safe instruction tuning, a recent development, requires more exploration. We hope that our work can shed light on the LLMs' potential to both bolster and jeopardize cybersecurity.
References in corpus (107)
- Training language models to follow instructions with human feedback
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- A Survey of Large Language Models
- Gender bias and stereotypes in Large Language Models
- BloombergGPT: A Large Language Model for Finance
- CCNet: Extracting High Quality Monolingual Datasets from Web Crawl Data
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- Language Models (Mostly) Know What They Know
- The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only
- Large Language Models for Code: Security Hardening and Adversarial Testing
- LIMA: Less Is More for Alignment
- Getting pwn'd by AI: Penetration Testing with Large Language Models
- MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots
- Eight Things to Know about Large Language Models
- Fake News in Sheep's Clothing: Robust Fake News Detection Against LLM-Empowered Style Attacks
- Adversarial Training for Large Neural Language Models
- Prompt Injection attack against LLM-integrated Applications
- Large Language Models Can Be Strong Differentially Private Learners
- Jailbroken: How Does LLM Safety Training Fail?
- LLM in the Shell: Generative Honeypots
- Decoding the Threat Landscape : ChatGPT, FraudGPT, and WormGPT in Social Engineering Attacks
- The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset
- CodeT: Code Generation with Generated Tests
- Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human Supervision
- LLM Lies: Hallucinations are not Bugs, but Features as Adversarial Examples
- Beyond the Safeguards: Exploring the Security Risks of ChatGPT
- No More Manual Tests? Evaluating and Improving ChatGPT for Unit Test Generation
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
- Physical Side-Channel Attacks on Embedded Neural Networks: A Survey
- PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts
- Prompt as Triggers for Backdoor Attack: Examining the Vulnerability in Language Models
- "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
- FuzzLLM: A Novel and Universal Fuzzing Framework for Proactively Discovering Jailbreak Vulnerabilities in Large Language Models
- Beyond Memorization: Violating Privacy Via Inference with Large Language Models
- Poisoning Language Models During Instruction Tuning
- Baseline Defenses for Adversarial Attacks Against Aligned Language Models
- RRHF: Rank Responses to Align Language Models with Human Feedback without tears
- Can LLM-Generated Misinformation Be Detected?
- Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks
- FATE-LLM: A Industrial Grade Federated Learning Framework for Large Language Models
- LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked
- SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
- How Should Pre-Trained Language Models Be Fine-Tuned Towards Adversarial Robustness?
- PentestGPT: An LLM-empowered Automatic Penetration Testing Tool
- Conversational Health Agents: A Personalized LLM-Powered Agent Framework
- Use of LLMs for Illicit Purposes: Threats, Prevention Measures, and Vulnerabilities
- From Text to MITRE Techniques: Exploring the Malicious Use of Large Language Models for Generating Cyber Attack Payloads
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
- Comparing Traditional and LLM-based Search for Consumer Choice: A Randomized Experiment
- ChatUniTest: A Framework for LLM-Based Test Generation
- Evaluation of ChatGPT Model for Vulnerability Detection
- Large Language Models are Edge-Case Fuzzers: Testing Deep Learning Libraries via FuzzGPT
- On the Reliability of Watermarks for Large Language Models
- Detecting Phishing Sites Using ChatGPT
- Low-code LLM: Graphical User Interface over Large Language Models
- Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
- Scaling Laws and Interpretability of Learning from Repeated Data
- GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
- Benchmarking Machine Translation with Cultural Awareness
- Chatbots to ChatGPT in a Cybersecurity Space: Evolution, Vulnerabilities, Attacks, Challenges, and Future Recommendations
- On The Impact of Machine Learning Randomness on Group Fairness
- Augmenting Greybox Fuzzing with Generative AI
- LLMs as Hackers: Autonomous Linux Privilege Escalation Attacks
- RatGPT: Turning online LLMs into Proxies for Malware Attacks
- Membership Inference on Word Embedding and Beyond
- Devising and Detecting Phishing: Large Language Models vs. Smaller Human Models
- Fake News Detectors are Biased against Texts Generated by Large Language Models
- How Prevalent is Gender Bias in ChatGPT? -- Exploring German and English ChatGPT Responses
- DIVAS: An LLM-based End-to-End Framework for SoC Security Analysis and Policy-based Protection
- On the Exploitability of Instruction Tuning
- Differentially Private Decoding in Large Language Models
- Pop Quiz! Can a Large Language Model Help With Reverse Engineering?
- Nightshade: Prompt-Specific Poisoning Attacks on Text-to-Image Generative Models
- Binary Code Summarization: Benchmarking ChatGPT/GPT-4 and Other Large Language Models
- Text Revealer: Private Text Reconstruction via Model Inversion Attacks against Transformers
- How Can ChatGPT Support Human Security Testers to Help Mitigate Supply Chain Attacks?
- Adversarial Attacks and Defenses in Large Language Models: Old and New Threats
- Weakly Supervised Veracity Classification with LLM-Predicted Credibility Signals
- Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions
- AutoDroid: LLM-powered Task Automation in Android
- Defending Pre-trained Language Models as Few-shot Learners against Backdoor Attacks
- Can Large Language Models Find And Fix Vulnerable Software?
- Identifying and Mitigating Privacy Risks Stemming from Language Models: A Survey
- Med-MMHL: A Multi-Modal Dataset for Detecting Human- and LLM-Generated Misinformation in the Medical Domain
- An Efficient Membership Inference Attack for the Diffusion Model by Proximal Initialization
- Shifting Attention to Relevance: Towards the Predictive Uncertainty Quantification of Free-Form Large Language Models
- Backdoor Attacks for In-Context Learning with Language Models
- Practical Membership Inference Attacks against Fine-tuned Large Language Models via Self-prompt Calibration
- REMARK-LLM: A Robust and Efficient Watermarking Framework for Generative Large Language Models
- Effective Prompt Extraction from Language Models
- Universal Jailbreak Backdoors from Poisoned Human Feedback
- Using Large Language Models for Cybersecurity Capture-The-Flag Challenges and Certification Questions
- DefectHunter: A Novel LLM-Driven Boosted-Conformer-based Code Vulnerability Detection Mechanism
- Probing Explicit and Implicit Gender Bias through LLM Conditional Text Generation
- Bot or Human? Detecting ChatGPT Imposters with A Single Question
- Revisiting Out-of-distribution Robustness in NLP: Benchmark, Analysis, and LLMs Evaluations
- Source Attribution for Large Language Model-Generated Data
- Prompt Packer: Deceiving LLMs through Compositional Instruction with Hidden Attacks
- Privacy Side Channels in Machine Learning Systems
- Using ChatGPT as a Static Application Security Testing Tool
- Self-Deception: Reverse Penetrating the Semantic Firewall of Large Language Models
- On the Safety of Open-Sourced Large Language Models: Does Alignment Really Prevent Them From Being Misused?
- LoFT: Local Proxy Fine-tuning For Improving Transferability Of Adversarial Attacks Against Large Language Model
- VulLibGen: Generating Names of Vulnerability-Affected Packages via a Large Language Model
- A Theoretical Insight into Attack and Defense of Gradient Leakage in Transformer
- A Probabilistic Fluctuation based Membership Inference Attack for Diffusion Models
- Test-time Backdoor Mitigation for Black-Box Large Language Models with Defensive Demonstrations
Cited by in corpus (51)
- Unleashing the potential of prompt engineering for large language models
- AI Agents Under Threat: A Survey of Key Security Challenges and Future Pathways
- Transformers and Large Language Models for Efficient Intrusion Detection Systems: A Comprehensive Survey
- Exploring the Capabilities and Limitations of Large Language Models in the Electric Energy Sector
- Tool Learning with Large Language Models: A Survey
- Survey of different Large Language Model Architectures: Trends, Benchmarks, and Challenges
- Generative AI for Self-Adaptive Systems: State of the Art and Research Roadmap
- Exploring the Roles of Large Language Models in Reshaping Transportation Systems: A Survey, Framework, and Roadmap
- Exploring the Potential of Large Language Models for Improving Digital Forensic Investigation Efficiency
- Exploring Gen-AI applications in building research and industry: A review
- An Empirical Study of Challenges in Machine Learning Asset Management
- An overview of domain-specific foundation model: key technologies, applications and challenges
- Mapping the Landscape of Generative AI in Network Monitoring and Management
- Embodied AI with Foundation Models for Mobile Service Robots: A Systematic Review
- LLMs as Hackers: Autonomous Linux Privilege Escalation Attacks
- SoK: The Privacy Paradox of Large Language Models: Advancements, Privacy Risks, and Mitigation
- VulScribeR: Exploring RAG-based Vulnerability Augmentation with LLMs
- Understanding Users' Security and Privacy Concerns and Attitudes Towards Conversational AI Platforms
- Learning from Data Streams: An Overview and Update
- From Prediction to Explanation: Multimodal, Explainable, and Interactive Deepfake Detection Framework for Non-Expert Users
- Stylometry recognizes human and LLM-generated texts in short samples
- Advances in Artificial Intelligence: A Review for the Creative Industries
- Crystalline Material Discovery in the Era of Artificial Intelligence
- Assessing the Effectiveness of LLMs in Android Application Vulnerability Analysis
- Robustness and Cybersecurity in the EU Artificial Intelligence Act
- LLM-Driven APT Detection for 6G Wireless Networks: A Systematic Review and Taxonomy
- Religious Bias Landscape in Language and Text-to-Image Models: Analysis, Detection, and Debiasing Strategies
- LLLMs: A Data-Driven Survey of Evolving Research on Limitations of Large Language Models
- PILLAR: an AI-Powered Privacy Threat Modeling Tool
- SecMLOps: A Comprehensive Framework for Integrating Security Throughout the MLOps Lifecycle
- Comparing the Performance of LLMs in RAG-based Question-Answering: A Case Study in Computer Science Literature
- Verifying the Robustness of Automatic Credibility Assessment
- Large Reasoning Models Are Autonomous Jailbreak Agents
- VFLAIR-LLM: A Comprehensive Framework and Benchmark for Split Learning of LLMs
- LLMs Have Rhythm: Fingerprinting Large Language Models Using Inter-Token Times and Network Traffic Analysis
- Using Large Language Models for Solving Thermodynamic Problems
- Integrating Artificial Open Generative Artificial Intelligence into Software Supply Chain Security
- Actionable Cybersecurity Notifications for Smart Homes: A User Study on the Role of Length and Complexity
- SAFE: Advancing Large Language Models in Leveraging Semantic and Syntactic Relationships for Software Vulnerability Detection
- LM-Fix: Lightweight Bit-Flip Detection and Rapid Recovery Framework for Language Models
- Transforming Remanufacturing Automation with Large Language Models: A Forward-Looking Analysis with Case Studies
- Federated Learning With Individualized Privacy Through Client Sampling
- Exploiting Latent Space Discontinuities for Building Universal LLM Jailbreaks and Data Extraction Attacks
- Context Conquers Parameters: Outperforming Proprietary LLM in Commit Message Generation
- REBot: From RAG to CatRAG with Semantic Enrichment and Graph Routing
- LLM-Augmented Knowledge Base Construction For Root Cause Analysis
- Exploring Potential Prompt Injection Attacks in Federated Military LLMs and Their Mitigation
- Mitigating Sensitive Information Leakage in LLMs4Code through Machine Unlearning
- Trustworthy AI LLM Scalability Risk Index (LSRI): A Cybersecurity Framework Assessing Agentic-AI Security & Software Model Supply Chain Safety Boosting AI-Generated Malware Defense & Explainability Mitigating Emerging Risks of Generative AI
- PRISM-XR: Empowering Privacy-Aware XR Collaboration with Multimodal Large Language Models
- Never say never: Exploring the effects of available knowledge on agent persuasiveness in controlled physiotherapy motivation dialogues