Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model
arXiv:1911.08265 · doi:10.1038/s41586-020-03051-4
Abstract
Constructing agents with planning capabilities has long been one of the main challenges in the pursuit of artificial intelligence. Tree-based planning methods have enjoyed huge success in challenging domains, such as chess and Go, where a perfect simulator is available. However, in real-world problems the dynamics governing the environment are often complex and unknown. In this work we present the MuZero algorithm which, by combining a tree-based search with a learned model, achieves superhuman performance in a range of challenging and visually complex domains, without any knowledge of their underlying dynamics. MuZero learns a model that, when applied iteratively, predicts the quantities most directly relevant to planning: the reward, the action-selection policy, and the value function. When evaluated on 57 different Atari games - the canonical video game environment for testing AI techniques, in which model-based planning approaches have historically struggled - our new algorithm achieved a new state of the art. When evaluated on Go, chess and shogi, without any knowledge of the game rules, MuZero matched the superhuman performance of the AlphaZero algorithm that was supplied with the game rules.
References in corpus (11)
- Learning to Plan Chemical Syntheses
- The Arcade Learning Environment: An Evaluation Platform for General Agents
- DeepStack: Expert-Level Artificial Intelligence in No-Limit Poker
- IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
- Massively Parallel Methods for Deep Reinforcement Learning
- Learning Continuous Control Policies by Stochastic Value Gradients
- Reinforcement Learning with Unsupervised Auxiliary Tasks
- OpenSpiel: A Framework for Reinforcement Learning in Games
- Observe and Look Further: Achieving Consistent Performance on Atari
- Learning and Querying Fast Generative Models for Reinforcement Learning
- Surprising Negative Results for Generative Adversarial Tree Search
Cited by in corpus (68)
- Efficient Deep Learning: A Survey on Making Deep Learning Models Smaller, Faster, and Better
- Physics Informed Neural Networks for Control Oriented Thermal Modeling of Buildings
- Partially Observable Markov Decision Processes in Robotics: A Survey
- Survey on reinforcement learning for language processing
- Survey on Large Language Model-Enhanced Reinforcement Learning: Concept, Taxonomy, and Methods
- Harms from Increasingly Agentic Algorithmic Systems
- Machine Culture
- General Purpose Artificial Intelligence Systems (GPAIS): Properties, Definition, Taxonomy, Societal Implications and Responsible Governance
- Data-Efficient Deep Reinforcement Learning for Attitude Control of Fixed-Wing UAVs: Field Experiments
- Materials Representation and Transfer Learning for Multi-Property Prediction
- A Survey on Reinforcement Learning in Aviation Applications
- Generative AI and Process Systems Engineering: The Next Frontier
- From Chess and Atari to StarCraft and Beyond: How Game AI is Driving the World of AI
- Deep reinforcement learning for machine scheduling: Methodology, the state-of-the-art, and future directions
- The Free Energy Principle for Perception and Action: A Deep Learning Perspective
- Visual Foresight Trees for Object Retrieval from Clutter with Nonprehensile Rearrangement
- Chess AI: Competing Paradigms for Machine Intelligence
- Optimal active particle navigation meets machine learning
- Parameterized Reinforcement Learning for Optical System Optimization
- Managing power grids through topology actions: A comparative study between advanced rule-based and reinforcement learning agents
- Rethinking Closed-loop Training for Autonomous Driving
- Sample-efficient Reinforcement Learning Representation Learning with Curiosity Contrastive Forward Dynamics Model
- GPT for Games: A Scoping Review (2020-2023)
- Multi-task convolutional neural network for image aesthetic assessment
- Interpretable pipelines with evolutionarily optimized modules for RL tasks with visual inputs
- Strategies for Using Proximal Policy Optimization in Mobile Puzzle Games
- Map-based Experience Replay: A Memory-Efficient Solution to Catastrophic Forgetting in Reinforcement Learning
- Beyond Games: A Systematic Review of Neural Monte Carlo Tree Search Applications
- Behavior Self-Organization Supports Task Inference for Continual Robot Learning
- Routing algorithms as tools for integrating social distancing with emergency evacuation
- Learning to design without prior data: Discovering generalizable design strategies using deep learning and tree search
- Reinforcement Learning with Dual-Observation for General Video Game Playing
- Reinforcement learning-based architecture search for quantum machine learning
- Towards Real-Time Routing Optimization with Deep Reinforcement Learning: Open Challenges
- Procedural Content Generation: Better Benchmarks for Transfer Reinforcement Learning
- Transforming Agency. On the mode of existence of Large Language Models
- A Review of Symbolic, Subsymbolic and Hybrid Methods for Sequential Decision Making
- Computational Performance of Deep Reinforcement Learning to find Nash Equilibria
- An Empirical Study on Google Research Football Multi-agent Scenarios
- Quantum State Reconstruction in a Noisy Environment via Deep Learning
- HUGO -- Highlighting Unseen Grid Options: Combining Deep Reinforcement Learning with a Heuristic Target Topology Approach
- Graph Reinforcement Learning for Power Grids: A Comprehensive Survey
- Weakly Supervised Disentangled Representation for Goal-conditioned Reinforcement Learning
- About optimal loss function for training physics-informed neural networks under respecting causality
- Mastering truss structure optimization with tree search
- Diversity-based Trajectory and Goal Selection with Hindsight Experience Replay
- The ProfessionAl Go annotation datasEt (PAGE)
- Centrally Coordinated Multi-Agent Reinforcement Learning for Power Grid Topology Control
- Impartial Games: A Challenge for Reinforcement Learning
- Combining AI Control Systems and Human Decision Support via Robustness and Criticality
- Intercepting Unauthorized Aerial Robots in Controlled Airspace Using Reinforcement Learning
- Simulated Mental Imagery for Robotic Task Planning
- Policy Gradients using Variational Quantum Circuits
- Tree-based machine learning performed in-memory with memristive analog CAM
- An Efficient Model-Agnostic Approach for Uncertainty Estimation in Data-Restricted Pedometric Applications
- Human Choice Prediction in Language-based Persuasion Games: Simulation-based Off-Policy Evaluation
- Generalist AI Control: Towards Multi-purpose Adaptive Algorithms
- Comparing Reinforcement Learning and Human Learning using the Game of Hidden Rules
- LuckyMera: a Modular AI Framework for Building Hybrid NetHack Agents
- Object-Centric World Models Meet Monte Carlo Tree Search
- UNSAT Solver Synthesis via Monte Carlo Forest Search
- Sample Efficient Reinforcement Learning via Large Vision Language Model Distillation
- Mobile Robot Exploration Without Maps via Out-of-Distribution Deep Reinforcement Learning
- Prime the search: Using large language models for guiding geometric task and motion planning by warm-starting tree search
- FORLORN: A Framework for Comparing Offline Methods and Reinforcement Learning for Optimization of RAN Parameters
- Augmenting the action space with conventions to improve multi-agent cooperation in Hanabi
- Learning disentangled representation for classical models
- Planning and Learning Using Adaptive Entropy Tree Search