StarCraft II: A New Challenge for Reinforcement Learning
arXiv:1708.04782
Abstract
This paper introduces SC2LE (StarCraft II Learning Environment), a reinforcement learning environment based on the StarCraft II game. This domain poses a new grand challenge for reinforcement learning, representing a more difficult class of problems than considered in most prior work. It is a multi-agent problem with multiple players interacting; there is imperfect information due to a partially observed map; it has a large action space involving the selection and control of hundreds of units; it has a large state space that must be observed solely from raw input feature planes; and it has delayed credit assignment requiring long-term strategies over thousands of steps. We describe the observation, action, and reward specification for the StarCraft II domain and provide an open source Python-based interface for communicating with the game engine. In addition to the main game maps, we provide a suite of mini-games focusing on different elements of StarCraft II gameplay. For the main game maps, we also provide an accompanying dataset of game replay data from human expert players. We give initial baseline results for neural networks trained from this data to predict game outcomes and player actions. Finally, we present initial baseline results for canonical deep reinforcement learning agents applied to the StarCraft II domain. On the mini-games, these agents learn to achieve a level of play that is comparable to a novice player. However, when trained on the main game, these agents are unable to make significant progress. Thus, SC2LE offers a new and challenging environment for exploring deep reinforcement learning algorithms and architectures.
Collaboration between DeepMind & Blizzard. 20 pages, 9 figures, 2 tables
References in corpus (3)
Cited by in corpus (70)
- A Brief Survey of Deep Reinforcement Learning
- Is Independent Learning All You Need in the StarCraft Multi-Agent Challenge?
- SMARTS: Scalable Multi-Agent Reinforcement Learning Training School for Autonomous Driving
- Supervised Learning Achieves Human-Level Performance in MOBA Games: A Case Study of Honor of Kings
- Dealing with Sparse Rewards in Reinforcement Learning
- Network Randomization: A Simple Technique for Generalization in Deep Reinforcement Learning
- MinAtar: An Atari-Inspired Testbed for Thorough and Reproducible Reinforcement Learning Experiments
- MAgent: A Many-Agent Reinforcement Learning Platform for Artificial Collective Intelligence
- Predicting Game Difficulty and Churn Without Players
- SEED RL: Scalable and Efficient Deep-RL with Accelerated Central Inference
- RLCard: A Toolkit for Reinforcement Learning in Card Games
- Autonomous Air Traffic Controller: A Deep Multi-Agent Reinforcement Learning Approach
- Network Environment Design for Autonomous Cyberdefense
- Auxiliary-task Based Deep Reinforcement Learning for Participant Selection Problem in Mobile Crowdsourcing
- Observational Overfitting in Reinforcement Learning
- Generation of ice states through deep reinforcement learning
- RMIX: Learning Risk-Sensitive Policies for Cooperative Reinforcement Learning Agents
- Evaluation of Human-AI Teams for Learned and Rule-Based Agents in Hanabi
- Heterogeneous Multi-Agent Reinforcement Learning for Unknown Environment Mapping
- Discovering Diverse Multi-Agent Strategic Behavior via Reward Randomization
- On Multi-Agent Learning in Team Sports Games
- Flatland-RL : Multi-Agent Reinforcement Learning on Trains
- Deep Reinforcement Learning with Pre-training for Time-efficient Training of Automatic Speech Recognition
- Cooperative Multi-Agent Transfer Learning with Level-Adaptive Credit Assignment
- On the Measure of Intelligence
- Using Fractal Neural Networks to Play SimCity 1 and Conway's Game of Life at Variable Scales
- Hybrid Actor-Critic Reinforcement Learning in Parameterized Action Space
- S2RMs: Spatially Structured Recurrent Modules
- Fever Basketball: A Complex, Flexible, and Asynchronized Sports Game Environment for Multi-agent Reinforcement Learning
- HALMA: Humanlike Abstraction Learning Meets Affordance in Rapid Problem Solving
- Using Unity to Help Solve Intelligence
- Do Autonomous Agents Benefit from Hearing?
- Neural Fictitious Self-Play on ELF Mini-RTS
- Pre-training in Deep Reinforcement Learning for Automatic Speech Recognition
- DeepCrawl: Deep Reinforcement Learning for Turn-based Strategy Games
- ToyBox: Better Atari Environments for Testing Reinforcement Learning Agents
- Lifelong Learning using Eigentasks: Task Separation, Skill Acquisition, and Selective Transfer
- A Narration-based Reward Shaping Approach using Grounded Natural Language Commands
- The Design Of "Stratega": A General Strategy Games Framework
- Human-in-the-Loop Methods for Data-Driven and Reinforcement Learning Systems
- StarCraft II Build Order Optimization using Deep Reinforcement Learning and Monte-Carlo Tree Search
- Arena: a toolkit for Multi-Agent Reinforcement Learning
- The Architectural Implications of Distributed Reinforcement Learning on CPU-GPU Systems
- Deep Latent Competition: Learning to Race Using Visual Control Policies in Latent Space
- On the notion of number in humans and machines
- Brittle AI, Causal Confusion, and Bad Mental Models: Challenges and Successes in the XAI Program
- Inferring Personalized Bayesian Embeddings for Learning from Heterogeneous Demonstration
- MimicBot: Combining Imitation and Reinforcement Learning to win in Bot Bowl
- Adversary agent reinforcement learning for pursuit-evasion
- A Brief Look at Generalization in Visual Meta-Reinforcement Learning
- Recurrent Value Functions
- Distributed Deep Reinforcement Learning: An Overview
- Dungeon Crawl Stone Soup as an Evaluation Domain for Artificial Intelligence
- Adaptive Agent Architecture for Real-time Human-Agent Teaming
- Stabilizing Q Learning Via Soft Mellowmax Operator
- Autonomous Exploration Under Uncertainty via Deep Reinforcement Learning on Graphs
- Exact Asymptotics for Linear Quadratic Adaptive Control
- Amortized Variational Deep Q Network
- First Draft on the xInf Model for Universal Physical Computation and Reverse Engineering of Natural Intelligence
- Growing Action Spaces
- DeepFoldit -- A Deep Reinforcement Learning Neural Network Folding Proteins
- State-based Episodic Memory for Multi-Agent Reinforcement Learning
- GrowSpace: Learning How to Shape Plants
- Applying supervised and reinforcement learning methods to create neural-network-based agents for playing StarCraft II
- Rethinking of AlphaStar
- Deep RL Agent for a Real-Time Action Strategy Game
- Autonomous Industrial Management via Reinforcement Learning: Self-Learning Agents for Decision-Making -- A Review
- Assured RL: Reinforcement Learning with Almost Sure Constraints
- DinerDash Gym: A Benchmark for Policy Learning in High-Dimensional Action Space
- Playing optical tweezers with deep reinforcement learning: in virtual, physical and augmented environments