Tactics of Adversarial Attack on Deep Reinforcement Learning Agents
arXiv:1703.06748
Abstract
We introduce two tactics to attack agents trained by deep reinforcement learning algorithms using adversarial examples, namely the strategically-timed attack and the enchanting attack. In the strategically-timed attack, the adversary aims at minimizing the agent's reward by only attacking the agent at a small subset of time steps in an episode. Limiting the attack activity to this subset helps prevent detection of the attack by the agent. We propose a novel method to determine when an adversarial example should be crafted and applied. In the enchanting attack, the adversary aims at luring the agent to a designated target state. This is achieved by combining a generative model and a planning algorithm: while the generative model predicts the future states, the planning algorithm generates a preferred sequence of actions for luring the agent. A sequence of adversarial examples is then crafted to lure the agent to take the preferred sequence of actions. We apply the two tactics to the agents trained by the state-of-the-art deep reinforcement learning algorithm including DQN and A3C. In 5 Atari games, our strategically timed attack reduces as much reward as the uniform attack (i.e., attacking at every time step) does by attacking the agent 4 times less often. Our enchanting attack lures the agent toward designated target states with a more than 70% success rate. Videos are available at http://yenchenlin.me/adversarial_attack_RL/
To Appear at IJCAI 2017. Project website: http://yenchenlin.me/adversarial_attack_RL/
Cited by in corpus (30)
- Certified Defenses for Data Poisoning Attacks
- Network Randomization: A Simple Technique for Generalization in Deep Reinforcement Learning
- Robust Reinforcement Learning on State Observations with Learned Optimal Adversary
- Adversarial Examples - A Complete Characterisation of the Phenomenon
- Detecting Adversarial Attacks on Neural Network Policies with Visual Foresight
- Security Matters: A Survey on Adversarial Machine Learning
- Adversarial Machine Learning: Bayesian Perspectives
- Action-Manipulation Attacks Against Stochastic Bandits: Attacks and Defense
- A Survey of Deep Reinforcement Learning in Recommender Systems: A Systematic Review and Future Directions
- Design of intentional backdoors in sequential models
- Reinforcement Learning with Perturbed Rewards
- On the Robustness of Cooperative Multi-Agent Reinforcement Learning
- Online Robustness Training for Deep Reinforcement Learning
- Snooping Attacks on Deep Reinforcement Learning
- Fooling Vision and Language Models Despite Localization and Attention Mechanism
- Adversarial Attacks on Reinforcement Learning based Energy Management Systems of Extended Range Electric Delivery Vehicles
- Optimal Attack and Defense for Reinforcement Learning
- Defense Against Reward Poisoning Attacks in Reinforcement Learning
- Evaluation of Momentum Diverse Input Iterative Fast Gradient Sign Method (M-DI2-FGSM) Based Attack Method on MCS 2018 Adversarial Attacks on Black Box Face Recognition System
- Generalization of Reinforcement Learning with Policy-Aware Adversarial Data Augmentation
- Adversarial Reinforcement Learning under Partial Observability in Autonomous Computer Network Defence
- Reward Poisoning in Reinforcement Learning: Attacks Against Unknown Learners in Unknown Environments
- Preventing Imitation Learning with Adversarial Policy Ensembles
- Robustness to Adversarial Attacks in Learning-Enabled Controllers
- Learning to Cope with Adversarial Attacks
- Minimalistic Attacks: How Little it Takes to Fool a Deep Reinforcement Learning Policy
- Deep Learning-Based Autonomous Driving Systems: A Survey of Attacks and Defenses
- Contributions to Large Scale Bayesian Inference and Adversarial Machine Learning
- Are You Tampering With My Data?
- Spatiotemporal Attacks for Embodied Agents