A Closer Look at Invalid Action Masking in Policy Gradient Algorithms
arXiv:2006.14171 · doi:10.32473/flairs.v35i.130584
Abstract
In recent years, Deep Reinforcement Learning (DRL) algorithms have achieved state-of-the-art performance in many challenging strategy games. Because these games have complicated rules, an action sampled from the full discrete action distribution predicted by the learned policy is likely to be invalid according to the game rules (e.g., walking into a wall). The usual approach to deal with this problem in policy gradient algorithms is to "mask out" invalid actions and just sample from the set of valid actions. The implications of this process, however, remain under-investigated. In this paper, we 1) show theoretical justification for such a practice, 2) empirically demonstrate its importance as the space of invalid actions grows, and 3) provide further insights by evaluating different action masking regimes, such as removing masking after an agent has been trained using masking. The source code can be found at https://github.com/vwxyzjn/invalid-action-masking
Accepted into the proceedings of International FLAIRS Conference Proceedings, Vol. 35 (2022)
References in corpus (1)
Cited by in corpus (28)
- CleanRL: High-quality Single-file Implementations of Deep Reinforcement Learning Algorithms
- Task Placement and Resource Allocation for Edge Machine Learning: A GNN-based Multi-Agent Reinforcement Learning Paradigm
- A Reinforcement Learning Environment For Job-Shop Scheduling
- Multi-agent Deep Reinforcement Learning for Distributed Load Restoration
- Provably Safe Reinforcement Learning via Action Projection using Reachability Analysis and Polynomial Zonotopes
- Deep Multi-agent Reinforcement Learning for Highway On-Ramp Merging in Mixed Traffic
- Frontier Semantic Exploration for Visual Target Navigation
- Introducing PetriRL: An Innovative Framework for JSSP Resolution Integrating Petri nets and Event-based Reinforcement Learning
- Generating Synergistic Formulaic Alpha Collections via Reinforcement Learning
- HOPE: A Reinforcement Learning-based Hybrid Policy Path Planner for Diverse Parking Scenarios
- Learning Hierarchical Interactive Multi-Object Search for Mobile Manipulation
- Learning to design without prior data: Discovering generalizable design strategies using deep learning and tree search
- Learning Based Dynamic Cluster Reconfiguration for UAV Mobility Management with 3D Beamforming
- Circuit Partitioning for Multi-Core Quantum Architectures with Deep Reinforcement Learning
- Egret: Reinforcement Mechanism for Sequential Computation Offloading in Edge Computing
- Understanding Continual Learning Settings with Data Distribution Drift Analysis
- Action Guidance: Getting the Best of Sparse Rewards and Shaped Rewards for Real-time Strategy Games
- Intent-based Radio Scheduler for RAN Slicing: Learning to deal with different network scenarios
- Novel Multi-Agent Action Masked Deep Reinforcement Learning for General Industrial Assembly Lines Balancing Problems
- Learning to solve arithmetic problems with a virtual abacus
- Minor Embedding for Quantum Annealing with Reinforcement Learning
- Learning to Generate All Feasible Actions
- Semi-on-Demand Transit Feeders with Shared Autonomous Vehicles and Reinforcement-Learning-Based Zonal Dispatching Control
- A Competition Winning Deep Reinforcement Learning Agent in microRTS
- RELiQ: Scalable Entanglement Routing via Reinforcement Learning in Quantum Networks
- Cooperative and Asynchronous Transformer-based Mission Planning for Heterogeneous Teams of Mobile Robots
- When Robots Say No: The Empathic Ethical Disobedience Benchmark
- Goal-Conditioned Reinforcement Learning for Data-Driven Maritime Navigation