Thinking Fast and Slow with Deep Learning and Tree Search
arXiv:1705.08439
Abstract
Sequential decision making problems, such as structured prediction, robotic control, and game playing, require a combination of planning policies and generalisation of those plans. In this paper, we present Expert Iteration (ExIt), a novel reinforcement learning algorithm which decomposes the problem into separate planning and generalisation tasks. Planning new policies is performed by tree search, while a deep neural network generalises those plans. Subsequently, tree search is improved by using the neural network policy to guide search, increasing the strength of new plans. In contrast, standard deep Reinforcement Learning algorithms rely on a neural network not only to generalise plans, but to discover them too. We show that ExIt outperforms REINFORCE for training a neural network to play the board game Hex, and our final tree search agent, trained tabula rasa, defeats MoHex 1.0, the most recent Olympiad Champion player to be publicly released.
v1 to v2: - Add a value function in MCTS - Some MCTS hyper-parameters changed - Repetition of experiments: improved accuracy and errors shown. (note the reduction in effect size for the tpt/cat experiment) - Results from a longer training run, including changes in expert strength in training - Comparison to MoHex. v3: clarify independence of ExIt and AG0. v4: see appendix E
References in corpus (2)
Cited by in corpus (47)
- Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
- Neo: A Learned Query Optimizer
- Scalable agent alignment via reward modeling: a research direction
- Balsa: Learning a Query Optimizer Without Expert Demonstrations
- Recursively Summarizing Books with Human Feedback
- Combining Deep Reinforcement Learning and Search for Imperfect-Information Games
- Plan Online, Learn Offline: Efficient Learning and Exploration via Model-Based Control
- Generative Language Modeling for Automated Theorem Proving
- Monte Carlo Tree Search: A Review of Recent Modifications and Applications
- Ranked Reward: Enabling Self-Play Reinforcement Learning for Combinatorial Optimization
- Guiding Deep Molecular Optimization with Genetic Exploration
- Learning to Search with MCTSnets
- Supervising strong learners by amplifying weak experts
- Learning Clustered Representation for Complex Free Energy Landscapes
- On-Policy Robot Imitation Learning from a Converging Supervisor
- Leveling the Playing Field -- Fairness in AI Versus Human Game Benchmarks
- Solving the Rubik's Cube Without Human Knowledge
- Goal-Directed Design Agents: Integrating Visual Imitation with One-Step Lookahead Optimization for Generative Design
- Learning to Play No-Press Diplomacy with Best Response Policy Iteration
- Learning to design without prior data: Discovering generalizable design strategies using deep learning and tree search
- Deep Reinforcement Learning Designed Shinnar-Le Roux RF Pulse using Root-Flipping: DeepRF_SLR
- Accelerating Cooperative Planning for Automated Vehicles with Learned Heuristics and Monte Carlo Tree Search
- Improving Neural Network Training using Dynamic Learning Rate Schedule for PINNs and Image Classification
- Scaling Scaling Laws with Board Games
- Adaptive Online Planning for Continual Lifelong Learning
- Using Monte Carlo Tree Search as a Demonstrator within Asynchronous Deep RL
- Deep Model-Based Reinforcement Learning for High-Dimensional Problems, a Survey
- Local Search for Policy Iteration in Continuous Control
- Information Theoretic Model Predictive Q-Learning
- Scalable Online Planning via Reinforcement Learning Fine-Tuning
- Synergising Human-like Responses and Machine Intelligence for Planning in Disaster Response
- Active Reinforcement Learning with Monte-Carlo Tree Search
- Transfer of Fully Convolutional Policy-Value Networks Between Games and Game Variants
- Learning Policies from Self-Play with Policy Gradients and MCTS Value Estimates
- Leveraging Statistical Multi-Agent Online Planning with Emergent Value Function Approximation
- Interleaving Fast and Slow Decision Making
- Few Shot System Identification for Reinforcement Learning
- Planning with a Receding Horizon for Manipulation in Clutter using a Learned Value Function
- Influence-Augmented Online Planning for Complex Environments
- Learning to Play Two-Player Perfect-Information Games without Knowledge
- Hedging of Financial Derivative Contracts via Monte Carlo Tree Search
- Self-Improved Retrosynthetic Planning
- Deep Pepper: Expert Iteration based Chess agent in the Reinforcement Learning Setting
- Single-Agent Optimization Through Policy Iteration Using Monte-Carlo Tree Search
- EXIT Analysis for Community Detection
- On the potential for open-endedness in neural networks
- Sidekick Policy Learning for Active Visual Exploration