Scaling Scaling Laws with Board Games
arXiv:2104.03113
Abstract
The largest experiments in machine learning now require resources far beyond the budget of all but a few institutions. Fortunately, it has recently been shown that the results of these huge experiments can often be extrapolated from the results of a sequence of far smaller, cheaper experiments. In this work, we show that not only can the extrapolation be done based on the size of the model, but on the size of the problem as well. By conducting a sequence of experiments using AlphaZero and Hex, we show that the performance achievable with a fixed amount of compute degrades predictably as the game gets larger and harder. Along with our main result, we further show that the test-time and train-time compute available to an agent can be traded off while maintaining performance.
References in corpus (10)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Scaling Laws for Neural Language Models
- Deep Learning Scaling is Predictable, Empirically
- Scaling Laws for Autoregressive Generative Modeling
- An Empirical Model of Large-Batch Training
- Phasic Policy Gradient
- Trivializations for Gradient-Based Optimization on Manifolds
- Monte-Carlo Tree Search as Regularized Policy Optimization
- Scaling Laws for Transfer
- Learning Curve Theory