CopyCAT: Taking Control of Neural Policies with Constant Attacks
arXiv:1905.12282
Abstract
We propose a new perspective on adversarial attacks against deep reinforcement learning agents. Our main contribution is CopyCAT, a targeted attack able to consistently lure an agent into following an outsider's policy. It is pre-computed, therefore fast inferred, and could thus be usable in a real-time scenario. We show its effectiveness on Atari 2600 games in the novel read-only setting. In this setting, the adversary cannot directly modify the agent's state -- its representation of the environment -- but can only attack the agent's observation -- its perception of the environment. Directly modifying the agent's state would require a write-access to the agent's inner workings and we argue that this assumption is too strong in realistic settings.
AAMAS 2020
References in corpus (7)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
- Rainbow: Combining Improvements in Deep Reinforcement Learning
- Robust Adversarial Reinforcement Learning
- Defensive Distillation is Not Robust to Adversarial Examples
- Dopamine: A Research Framework for Deep Reinforcement Learning
- Delving into adversarial attacks on deep policies
Cited by in corpus (5)
- When and How to Fool Explainable Models (and Humans) with Adversarial Examples
- Towards Resilient Artificial Intelligence: Survey and Research Issues
- Adversarial Attacks on Linear Contextual Bandits
- Stealthy and Efficient Adversarial Attacks against Deep Reinforcement Learning
- CROP: Certifying Robust Policies for Reinforcement Learning through Functional Smoothing