AI Agents Under Threat: A Survey of Key Security Challenges and Future Pathways
arXiv:2406.02630 · doi:10.1145/3716628
Abstract
An Artificial Intelligence (AI) agent is a software entity that autonomously performs tasks or makes decisions based on pre-defined objectives and data inputs. AI agents, capable of perceiving user inputs, reasoning and planning tasks, and executing actions, have seen remarkable advancements in algorithm development and task performance. However, the security challenges they pose remain under-explored and unresolved. This survey delves into the emerging security threats faced by AI agents, categorizing them into four critical knowledge gaps: unpredictability of multi-step user inputs, complexity in internal executions, variability of operational environments, and interactions with untrusted external entities. By systematically reviewing these threats, this paper highlights both the progress made and the existing limitations in safeguarding AI agents. The insights provided aim to inspire further research into addressing the security threats associated with AI agents, thereby fostering the development of more robust and secure AI agent applications.
Submitted to ACM Computing Survey
References in corpus (9)
- Training language models to follow instructions with human feedback
- A Survey on Large Language Model (LLM) Security and Privacy: The Good, the Bad, and the Ugly
- Toolformer: Language Models Can Teach Themselves to Use Tools
- MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots
- From ChatGPT to ThreatGPT: Impact of Generative AI in Cybersecurity and Privacy
- FuzzLLM: A Novel and Universal Fuzzing Framework for Proactively Discovering Jailbreak Vulnerabilities in Large Language Models
- Toxicity Detection with Generative Prompt-based Inference
- Efficient Adversarial Attacks on Online Multi-agent Reinforcement Learning
- Robust Multi-Agent Reinforcement Learning via Adversarial Regularization: Theoretical Foundation and Stable Algorithms