Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
arXiv:1712.01815
Abstract
The game of chess is the most widely-studied domain in the history of artificial intelligence. The strongest programs are based on a combination of sophisticated search techniques, domain-specific adaptations, and handcrafted evaluation functions that have been refined by human experts over several decades. In contrast, the AlphaGo Zero program recently achieved superhuman performance in the game of Go, by tabula rasa reinforcement learning from games of self-play. In this paper, we generalise this approach into a single AlphaZero algorithm that can achieve, tabula rasa, superhuman performance in many challenging domains. Starting from random play, and given no domain knowledge except the game rules, AlphaZero achieved within 24 hours a superhuman level of play in the games of chess and shogi (Japanese chess) as well as Go, and convincingly defeated a world-champion program in each case.
References in corpus (2)
Cited by in corpus (112)
- Solving Rubik's Cube with a Robot Hand
- Artificial Intelligence and Big Data in Entrepreneurship: A New Era Has Begun
- Ablation Studies in Artificial Neural Networks
- Building a Conversational Agent Overnight with Dialogue Self-Play
- Dealing with Non-Stationarity in Multi-Agent Deep Reinforcement Learning
- Harms from Increasingly Agentic Algorithmic Systems
- SMARTS: Scalable Multi-Agent Reinforcement Learning Training School for Autonomous Driving
- Innateness, AlphaZero, and Artificial Intelligence
- Digital Twin: Values, Challenges and Enablers
- Neurosymbolic Reinforcement Learning and Planning: A Survey
- On the Binding Problem in Artificial Neural Networks
- The Free Energy Principle for Perception and Action: A Deep Learning Perspective
- Generative Language Modeling for Automated Theorem Proving
- Making Deep Q-learning methods robust to time discretization
- Deep Learning Based Simulators for the Phosphorus Removal Process Control in Wastewater Treatment via Deep Reinforcement Learning Algorithms
- How Much Automation Does a Data Scientist Want?
- Reassessing Claims of Human Parity and Super-Human Performance in Machine Translation at WMT 2019
- The Chess Transformer: Mastering Play using Generative Language Models
- AlphaD3M: Machine Learning Pipeline Synthesis
- Asymmetric self-play for automatic goal discovery in robotic manipulation
- Policy Gradient Search: Online Planning and Expert Iteration without Search Trees
- Open Problems in Cooperative AI
- Hindsight Goal Ranking on Replay Buffer for Sparse Reward Environment
- Modern Deep Reinforcement Learning Algorithms
- Reframing Jet Physics with New Computational Methods
- Reinforcement Learning-Driven Test Generation for Android GUI Applications using Formal Specifications
- On Multi-Agent Learning in Team Sports Games
- Solving Rubik's Cube via Quantum Mechanics and Deep Reinforcement Learning
- Robust Reinforcement Learning in POMDPs with Incomplete and Noisy Observations
- Affordance-based Reinforcement Learning for Urban Driving
- Hyper-Parameter Sweep on AlphaZero General
- Learning hierarchical behavior and motion planning for autonomous driving
- Machine Learning Challenges and Opportunities in the African Agricultural Sector -- A General Perspective
- Scaling Imitation Learning in Minecraft
- Stochastic reserving with a stacked model based on a hybridized Artificial Neural Network
- Learning to Communicate in Multi-Agent Reinforcement Learning : A Review
- A Versatile Multi-Robot Monte Carlo Tree Search Planner for On-Line Coverage Path Planning
- Continuous Control with Deep Reinforcement Learning for Autonomous Vessels
- Explainable AI: Deep Reinforcement Learning Agents for Residential Demand Side Cost Savings in Smart Grids
- Solving Hard AI Planning Instances Using Curriculum-Driven Deep Reinforcement Learning
- Neural Fictitious Self-Play on ELF Mini-RTS
- Adaptive Monte Carlo Multiple Testing via Multi-Armed Bandits
- Rethinking Cooperative Rationalization: Introspective Extraction and Complement Control
- State-of-Charge Estimation of a Li-Ion Battery using Deep Forward Neural Networks
- Learning Agile Locomotion via Adversarial Training
- Catalyst.RL: A Distributed Framework for Reproducible RL Research
- Skin disease diagnosis with deep learning: a review
- StarCraft II Build Order Optimization using Deep Reinforcement Learning and Monte-Carlo Tree Search
- Convergence of Multi-Agent Learning with a Finite Step Size in General-Sum Games
- Self-play Learning Strategies for Resource Assignment in Open-RAN Networks
- Train on Small, Play the Large: Scaling Up Board Games with AlphaZero and GNN
- RL STaR Platform: Reinforcement Learning for Simulation based Training of Robots
- The Adversarial Resilience Learning Architecture for AI-based Modelling, Exploration, and Operation of Complex Cyber-Physical Systems
- Graceful Degradation and Related Fields
- Playing Chess with Limited Look Ahead
- Competitive Experience Replay
- Towards Understanding Chinese Checkers with Heuristics, Monte Carlo Tree Search, and Deep Reinforcement Learning
- Runtime Verification of Learning Properties for Reinforcement Learning Algorithms
- On Reinforcement Learning for Turn-based Zero-sum Markov Games
- On Learning to Prove
- Multi-Agent Deep Reinforcement Learning with Adaptive Policies
- Towards Automatic Actor-Critic Solutions to Continuous Control
- Can You Learn an Algorithm? Generalizing from Easy to Hard Problems with Recurrent Networks
- Planning from Pixels using Inverse Dynamics Models
- LiveChess2FEN: a Framework for Classifying Chess Pieces based on CNNs
- Continuous Control for Searching and Planning with a Learned Model
- Blending MPC & Value Function Approximation for Efficient Reinforcement Learning
- Sample Efficient Ensemble Learning with Catalyst.RL
- Hierarchical clustering in particle physics through reinforcement learning
- Reducing catastrophic forgetting when evolving neural networks
- Typed Graph Networks
- Fractional Transfer Learning for Deep Model-Based Reinforcement Learning
- Model-based Meta Reinforcement Learning using Graph Structured Surrogate Models
- Ignorance-Aware Approaches and Algorithms for Prototype Selection in Machine Learning
- The Matthew Effect in Computation Contests: High Difficulty May Lead to 51% Dominance
- Utility Ghost: Gamified redistricting with partisan symmetry
- Multi-Issue Bargaining With Deep Reinforcement Learning
- Finite Group Equivariant Neural Networks for Games
- Time Adaptive Reinforcement Learning
- How Do You Act? An Empirical Study to Understand Behavior of Deep Reinforcement Learning Agents
- Mastering Terra Mystica: Applying Self-Play to Multi-agent Cooperative Board Games
- Scaffolding Reflection in Reinforcement Learning Framework for Confinement Escape Problem
- Augmented Shortcuts for Vision Transformers
- Learning proofs for the classification of nilpotent semigroups
- Mack-Net model: Blending Mack's model with Recurrent Neural Networks
- BridgeHand2Vec Bridge Hand Representation
- Comparison Training for Computer Chinese Chess
- Analyzing the Hidden Activations of Deep Policy Networks: Why Representation Matters
- Deep Synoptic Monte Carlo Planning in Reconnaissance Blind Chess
- Neural Networks and Denotation
- Learning Cooperation and Online Planning Through Simulation and Graph Convolutional Network
- Deep Reinforcement Learning with Adjustments
- On the Power of Refined Skat Selection
- Searching with Opponent-Awareness
- Pyfectious: An individual-level simulator to discover optimal containment polices for epidemic diseases
- Evolution of Q Values for Deep Q Learning in Stable Baselines
- Learning to Play Foosball: System and Baselines
- Marginal Utility for Planning in Continuous or Large Discrete Action Spaces
- Safety Aware Reinforcement Learning (SARL)
- Growing Action Spaces
- Automated Chess Commentator Powered by Neural Chess Engine
- Autonomous Industrial Management via Reinforcement Learning: Self-Learning Agents for Decision-Making -- A Review
- Room Clearance with Feudal Hierarchical Reinforcement Learning
- Self-Play Learning Without a Reward Metric
- Near-optimal Bayesian Solution For Unknown Discrete Markov Decision Process
- Deep RL Agent for a Real-Time Action Strategy Game
- Neural-Network Guided Expression Transformation
- Solving QSAT problems with neural MCTS
- Detection of chaotic behavior in time series
- Enhanced Scene Specificity with Sparse Dynamic Value Estimation
- NEARL: Non-Explicit Action Reinforcement Learning for Robotic Control
- Network Horizon Dynamics I: Qualitative Aspects