Finding Effective Security Strategies through Reinforcement Learning and Self-Play
arXiv:2009.08120 · doi:10.23919/CNSM50824.2020.9269092
Abstract
We present a method to automatically find security strategies for the use case of intrusion prevention. Following this method, we model the interaction between an attacker and a defender as a Markov game and let attack and defense strategies evolve through reinforcement learning and self-play without human intervention. Using a simple infrastructure configuration, we demonstrate that effective security strategies can emerge from self-play. This shows that self-play, which has been applied in other domains with great success, can be effective in the context of network security. Inspection of the converged policies show that the emerged policies reflect common-sense knowledge and are similar to strategies of humans. Moreover, we address known challenges of reinforcement learning in this domain and present an approach that uses function approximation, an opponent pool, and an autoregressive policy representation. Through evaluations we show that our method is superior to two baseline methods but that policy convergence in self-play remains a challenge.
References in corpus (1)
Cited by in corpus (9)
- Learning Near-Optimal Intrusion Responses Against Dynamic Attackers
- Intrusion Prevention through Optimal Stopping
- Prospective Artificial Intelligence Approaches for Active Cyber Defence
- Learning Intrusion Prevention Policies through Optimal Stopping
- Out of the Cage: How Stochastic Parrots Win in Cyber Security Environments
- Automated Security Response through Online Learning with Adaptive Conjectures
- Canaries and Whistles: Resilient Drone Communication Networks with (or without) Deep Reinforcement Learning
- Catch Me If You Can: Improving Adversaries in Cyber-Security With Q-Learning Algorithms
- A System for Interactive Examination of Learned Security Policies