BACKDOORL: Backdoor Attack against Competitive Reinforcement Learning
arXiv:2105.00579
Abstract
Recent research has confirmed the feasibility of backdoor attacks in deep reinforcement learning (RL) systems. However, the existing attacks require the ability to arbitrarily modify an agent's observation, constraining the application scope to simple RL systems such as Atari games. In this paper, we migrate backdoor attacks to more complex RL systems involving multiple agents and explore the possibility of triggering the backdoor without directly manipulating the agent's observation. As a proof of concept, we demonstrate that an adversary agent can trigger the backdoor of the victim agent with its own action in two-player competitive RL systems. We prototype and evaluate BACKDOORL in four competitive environments. The results show that when the backdoor is activated, the winning rate of the victim drops by 17% to 37% compared to when not activated.
References in corpus (10)
- Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning
- Can You Really Backdoor Federated Learning?
- Safe, Multi-Agent, Reinforcement Learning for Autonomous Driving
- CARLA: An Open Urban Driving Simulator
- TABOR: A Highly Accurate Approach to Inspecting and Restoring Trojan Backdoors in AI Systems
- Gotta Catch 'Em All: Using Honeypots to Catch Adversarial Attacks on Neural Networks
- Label-Consistent Backdoor Attacks
- BAAAN: Backdoor Attacks Against Autoencoder and GAN-Based Machine Learning Models
- Design of intentional backdoors in sequential models
- Don't Trigger Me! A Triggerless Backdoor Attack Against Deep Neural Networks