Intrusion Prevention through Optimal Stopping
arXiv:2111.00289 · doi:10.1109/TNSM.2022.3176781
Abstract
We study automated intrusion prevention using reinforcement learning. Following a novel approach, we formulate the problem of intrusion prevention as an (optimal) multiple stopping problem. This formulation gives us insight into the structure of optimal policies, which we show to have threshold properties. For most practical cases, it is not feasible to obtain an optimal defender policy using dynamic programming. We therefore develop a reinforcement learning approach to approximate an optimal threshold policy. We introduce T-SPSA, an efficient reinforcement learning algorithm that learns threshold policies through stochastic approximation. We show that T-SPSA outperforms state-of-the-art algorithms for our use case. Our overall method for learning and validating policies includes two systems: a simulation system where defender policies are incrementally learned and an emulation system where statistics are produced that drive simulation runs and where learned policies are evaluated. We show that this approach can produce effective defender policies for a practical IT infrastructure.
Preprint; Submitted to IEEE for review. major revision 1/4 2022. arXiv admin note: substantial text overlap with arXiv:2106.07160
References in corpus (7)
- Selling a stock at the ultimate maximum
- Network Environment Design for Autonomous Cyberdefense
- Prospective Artificial Intelligence Approaches for Active Cyber Defence
- Deep hierarchical reinforcement agents for automated penetration testing
- Learning Intrusion Prevention Policies through Optimal Stopping
- CybORG: A Gym for the Development of Autonomous Cyber Agents
- Reinforcement Learning for Feedback-Enabled Cyber Resilience