Scaling Laws for Autoregressive Generative Modeling
arXiv:2010.14701
Abstract
We identify empirical scaling laws for the cross-entropy loss in four domains: generative image modeling, video modeling, multimodal imagetext models, and mathematical problem solving. In all cases autoregressive Transformers smoothly improve in performance as model size and compute budgets increase, following a power-law plus constant scaling law. The optimal model size also depends on the compute budget through a power-law, with exponents that are nearly universal across all data domains. The cross-entropy loss has an information theoretic interpretation as TrueTrueModel, and the empirical scaling laws suggest a prediction for both the true data distribution's entropy and the KL divergence between the true and model distributions. With this interpretation, billion-parameter Transformers are nearly perfect models of the YFCC100M image distribution downsampled to an resolution, and we can forecast the model size needed to achieve any given reducible loss (ie ) in nats/image for other resolutions. We find a number of additional scaling laws in specific domains: (a) we identify a scaling relation for the mutual information between captions and images in multimodal models, and show how to answer the question "Is a picture worth a thousand words?"; (b) in the case of mathematical problem solving, we identify scaling laws for model performance when extrapolating beyond the training distribution; (c) we finetune generative image models for ImageNet classification and find smooth scaling of the classification loss and error rate, even as the generative loss levels off. Taken together, these results strengthen the case that scaling laws have important implications for neural network performance, including on downstream tasks.
20+17 pages, 33 figures; added appendix with additional language results
References in corpus (10)
- Scaling Laws for Neural Language Models
- Generating Long Sequences with Sparse Transformers
- A Downsampled Variant of ImageNet as an Alternative to the CIFAR datasets
- An Empirical Model of Large-Batch Training
- Jukebox: A Generative Model for Music
- Analysing Mathematical Reasoning Abilities of Neural Models
- Generating Wikipedia by Summarizing Long Sequences
- The large learning rate phase of deep learning: the catapult mechanism
- A Neural Scaling Law from the Dimension of the Data Manifold
- One Epoch Is All You Need
Cited by in corpus (12)
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
- Revisiting ResNets: Improved Training and Scaling Strategies
- Investigating the Limitations of Transformers with Simple Arithmetic Tasks
- Scaling Laws for Transfer
- Generalization bounds for deep learning
- Unsupervised Neural Machine Translation with Generative Language Models Only
- Scaling Scaling Laws with Board Games
- Learning Curve Theory
- Universal Policies for Software-Defined MDPs
- Turing-Universal Learners with Optimal Scaling Laws
- Scaling Laws for the Few-Shot Adaptation of Pre-trained Image Classifiers