Acquisition of Chess Knowledge in AlphaZero
arXiv:2111.09259 · doi:10.1073/pnas.2206625119
Abstract
What is learned by sophisticated neural network agents such as AlphaZero? This question is of both scientific and practical interest. If the representations of strong neural networks bear no resemblance to human concepts, our ability to understand faithful explanations of their decisions will be restricted, ultimately limiting what we can achieve with neural network interpretability. In this work we provide evidence that human knowledge is acquired by the AlphaZero neural network as it trains on the game of chess. By probing for a broad range of human chess concepts we show when and where these concepts are represented in the AlphaZero network. We also provide a behavioural analysis focusing on opening play, including qualitative analysis from chess Grandmaster Vladimir Kramnik. Finally, we carry out a preliminary investigation looking at the low-level details of AlphaZero's representations, and make the resulting behavioural and representational analyses available online.
69 pages, 44 figures
References in corpus (15)
- Towards A Rigorous Science of Interpretable Machine Learning
- Understanding the Role of Individual Units in a Deep Neural Network
- Hybrid Reward Architecture for Reinforcement Learning
- The (Un)reliability of saliency methods
- Towards Deep Symbolic Reinforcement Learning
- Acquisition of Chess Knowledge in AlphaZero
- Aligning Superhuman AI with Human Behavior: Chess as a Model System
- DisCoRL: Continual Reinforcement Learning via Policy Distillation
- A Theory of Usable Information Under Computational Constraints
- Decoupling feature extraction from policy learning: assessing benefits of state representation learning in goal based robotics
- Concept-based model explanations for Electronic Health Records
- Promises and Pitfalls of Black-Box Concept Learning Models
- Assessing Game Balance with AlphaZero: Exploring Alternative Rule Sets in Chess
- DREAM Architecture: a Developmental Approach to Open-Ended Learning in Robotics
- Domain-Level Explainability -- A Challenge for Creating Trust in Superhuman AI Strategies
Cited by in corpus (6)
- Adversarial attacks and defenses in explainable artificial intelligence: A survey
- Acquisition of Chess Knowledge in AlphaZero
- Human-like object concept representations emerge naturally in multimodal large language models
- Learning Chess With Language Models and Transformers
- INSPECT: Intrinsic and Systematic Probing Evaluation for Code Transformers
- Birds look like cars: Adversarial analysis of intrinsically interpretable deep learning