Delving into adversarial attacks on deep policies
arXiv:1705.06452
Abstract
Adversarial examples have been shown to exist for a variety of deep learning architectures. Deep reinforcement learning has shown promising results on training agent policies directly on raw inputs such as image pixels. In this paper we present a novel study into adversarial attacks on deep reinforcement learning polices. We compare the effectiveness of the attacks using adversarial examples vs. random noise. We present a novel method for reducing the number of times adversarial examples need to be injected for a successful attack, based on the value function. We further explore how re-training on random noise and FGSM perturbations affects the resilience against adversarial examples.
ICLR 2017 Workshop
Cited by in corpus (5)
- Robust Deep Reinforcement Learning with Adversarial Attacks
- Whatever Does Not Kill Deep Reinforcement Learning, Makes It Stronger
- Detecting Adversarial Attacks on Neural Network Policies with Visual Foresight
- Better the Devil you Know: An Analysis of Evasion Attacks using Out-of-Distribution Adversarial Examples
- Augmenting Model Robustness with Transformation-Invariant Attacks