The Emerged Security and Privacy of LLM Agent: A Survey with Case Studies
arXiv:2407.19354 · doi:10.1145/3773080
Abstract
Inspired by the rapid development of Large Language Models (LLMs), LLM agents have evolved to perform complex tasks. LLM agents are now extensively applied across various domains, handling vast amounts of data to interact with humans and execute tasks. The widespread applications of LLM agents demonstrate their significant commercial value; however, they also expose security and privacy vulnerabilities. At the current stage, comprehensive research on the security and privacy of LLM agents is highly needed. This survey aims to provide a comprehensive overview of the newly emerged privacy and security issues faced by LLM agents. We begin by introducing the fundamental knowledge of LLM agents, followed by a categorization and analysis of the threats. We then discuss the impacts of these threats on humans, environment, and other agents. Subsequently, we review existing defensive strategies, and finally explore future trends. Additionally, the survey incorporates diverse case studies to facilitate a more accessible understanding. By highlighting these critical security and privacy issues, the survey seeks to stimulate future research towards enhancing the security and privacy of LLM agents, thereby increasing their reliability and trustworthiness in future applications.
35 pages, 19 figures. Accepted to ACM Computing Surveys (CSUR), 2025
References in corpus (70)
- ReAct: Synergizing Reasoning and Acting in Language Models
- A Survey on ChatGPT: AI-Generated Contents, Challenges, and Solutions
- The Rise and Potential of Large Language Model Based Agents: A Survey
- Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
- MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework
- ChemCrow: Augmenting large-language models with chemistry tools
- "It's a Fair Game", or Is It? Examining How Users Navigate Disclosure Risks and Benefits When Using LLM-Based Conversational Agents
- Prompt Injection attack against LLM-integrated Applications
- MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action
- Frontier AI Regulation: Managing Emerging Risks to Public Safety
- Large Language Model based Multi-Agents: A Survey of Progress and Challenges
- Decoding the Threat Landscape : ChatGPT, FraudGPT, and WormGPT in Social Engineering Attacks
- AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
- AgentBench: Evaluating LLMs as Agents
- Aligning Large Language Models with Human: A Survey
- Multi-Agent Collaboration: Harnessing the Power of Intelligent LLM Agents
- Chain-of-Verification Reduces Hallucination in Large Language Models
- Poisoning Language Models During Instruction Tuning
- Security and Privacy Challenges of Large Language Models: A Survey
- SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
- From Text to MITRE Techniques: Exploring the Malicious Use of Large Language Models for Generating Cyber Attack Payloads
- Toxicity in ChatGPT: Analyzing Persona-assigned Language Models
- Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
- Investigating the Catastrophic Forgetting in Multimodal Large Language Models
- An Embodied Generalist Agent in 3D World
- Habitat 3.0: A Co-Habitat for Humans, Avatars and Robots
- LLM Agents can Autonomously Exploit One-day Vulnerabilities
- AgentSims: An Open-Source Sandbox for Large Language Model Evaluation
- Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
- Aligning Large Multimodal Models with Factually Augmented RLHF
- Large Multimodal Agents: A Survey
- TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
- RoCo: Dialectic Multi-Robot Collaboration with Large Language Models
- Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language Models
- Zero-Resource Hallucination Prevention for Large Language Models
- Practical Membership Inference Attacks against Fine-tuned Large Language Models via Self-prompt Calibration
- Exposing Attention Glitches with Flip-Flop Language Modeling
- Characterizing Attribution and Fluency Tradeoffs for Retrieval-Augmented Large Language Models
- Unique Security and Privacy Threats of Large Language Models: A Comprehensive Survey
- Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based Agents
- VLATTACK: Multimodal Adversarial Attacks on Vision-Language Tasks via Pre-trained Models
- Weight Averaging: A Simple Yet Effective Method to Overcome Catastrophic Forgetting in Automatic Speech Recognition
- ParaFuzz: An Interpretability-Driven Technique for Detecting Poisoned Samples in NLP
- Enhancing Language Representation with Constructional Information for Natural Language Understanding
- Adapting LLM Agents with Universal Feedback in Communication
- InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents
- Identifying the Risks of LM Agents with an LM-Emulated Sandbox
- A Comprehensive Overview of Backdoor Attacks in Large Language Models within Communication Networks
- Defending Against Indirect Prompt Injection Attacks With Spotlighting
- Avalon's Game of Thoughts: Battle Against Deception through Recursive Contemplation
- Empowering Language Models with Active Inquiry for Deeper Understanding
- WALL-E: Embodied Robotic WAiter Load Lifting with Large Language Model
- Combing for Credentials: Active Pattern Extraction from Smart Reply
- Mitigating Catastrophic Forgetting in Large Language Models with Self-Synthesized Rehearsal
- Synergizing Human-AI Agency: A Guide of 23 Heuristics for Service Co-Creation with LLM-Based Agents
- Multi-Agent Collaboration via Cross-Team Orchestration
- Towards Action Hijacking of Large Language Model-based Agent
- Agent-SafetyBench: Evaluating the Safety of LLM Agents
- InferDPT: Privacy-Preserving Inference for Closed-box Large Language Model
- SAUP: Situation Awareness Uncertainty Propagation on LLM Agent
- MAgIC: Investigation of Large Language Model Powered Multi-Agent in Cognition, Adaptability, Rationality and Collaboration
- SafeAgentBench: A Benchmark for Safe Task Planning of Embodied LLM Agents
- Imprompter: Tricking LLM Agents into Improper Tool Use
- User Inference Attacks on Large Language Models
- Fine-tuning Large Language Models with Sequential Instructions
- HallE-Control: Controlling Object Hallucination in Large Multimodal Models
- CoCA: Regaining Safety-awareness of Multimodal Large Language Models with Constitutional Calibration
- Volcano: Mitigating Multimodal Hallucination through Self-Feedback Guided Revision
- LeakAgent: RL-based Red-teaming Agent for LLM Privacy Leakage
- V-IRL: Grounding Virtual Intelligence in Real Life