Training Compute-Optimal Large Language Models
arXiv:2203.15556
Abstract
We investigate the optimal model size and number of tokens for training a transformer language model under a given compute budget. We find that current large language models are significantly undertrained, a consequence of the recent focus on scaling language models whilst keeping the amount of training data constant. By training over 400 language models ranging from 70 million to over 16 billion parameters on 5 to 500 billion tokens, we find that for compute-optimal training, the model size and the number of training tokens should be scaled equally: for every doubling of model size the number of training tokens should also be doubled. We test this hypothesis by training a predicted compute-optimal model, Chinchilla, that uses the same compute budget as Gopher but with 70B parameters and 4 more more data. Chinchilla uniformly and significantly outperforms Gopher (280B), GPT-3 (175B), Jurassic-1 (178B), and Megatron-Turing NLG (530B) on a large range of downstream evaluation tasks. This also means that Chinchilla uses substantially less compute for fine-tuning and inference, greatly facilitating downstream usage. As a highlight, Chinchilla reaches a state-of-the-art average accuracy of 67.5% on the MMLU benchmark, greater than a 7% improvement over Gopher.
Cited by in corpus (51)
- DINOv2: Learning Robust Visual Features without Supervision
- Reproducible scaling laws for contrastive language-image learning
- Large Language Models (LLMs) as Agents for Augmented Democracy
- Compute Trends Across Three Eras of Machine Learning
- Large AI Models in Health Informatics: Applications, Challenges, and the Future
- Auditing large language models: a three-layered approach
- 14 Examples of How LLMs Can Transform Materials Science and Chemistry: A Reflection on a Large Language Model Hackathon
- What Large Language Models Know and What People Think They Know
- Large Language Models and the Reverse Turing Test
- Large Language Models and User Trust: Consequence of Self-Referential Learning Loop and the Deskilling of Healthcare Professionals
- Self-supervised Learning for Human Activity Recognition Using 700,000 Person-days of Wearable Data
- Machine Culture
- NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing
- ChatGPT Empowered Long-Step Robot Control in Various Environments: A Case Application
- The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset
- Neural Natural Language Processing for Long Texts: A Survey on Classification and Summarization
- A Comprehensive Survey on Integrating Large Language Models with Knowledge-Based Methods
- Evaluation of LLM Chatbots for OSINT-based Cyber Threat Awareness
- Enough With "Human-AI Collaboration"
- Data Governance in the Age of Large-Scale Data-Driven Language Technology
- A Survey on Symbolic Knowledge Distillation of Large Language Models
- A social path to human-like artificial intelligence
- Trends in Energy Estimates for Computing in AI/Machine Learning Accelerators, Supercomputers, and Compute-Intensive Applications
- Several categories of Large Language Models (LLMs): A Short Survey
- Advances in machine-learning-based sampling motivated by lattice quantum chromodynamics
- Transformer Grammars: Augmenting Transformer Language Models with Syntactic Inductive Biases at Scale
- How quantum computing can enhance biomarker discovery
- A Comprehensive Review of State-of-The-Art Methods for Java Code Generation from Natural Language Text
- System 2 thinking in OpenAI's o1-preview model: Near-perfect performance on a mathematics exam
- Fly-Swat or Cannon? Cost-Effective Language Model Choice via Meta-Modeling
- ReCycle: Resilient Training of Large DNNs using Pipeline Adaptation
- Large Language Models at Work in China's Labor Market
- Exploring the Landscape of Ubiquitous In-home Health Monitoring: A Comprehensive Survey
- Evidence of interrelated cognitive-like capabilities in large language models: Indications of artificial general intelligence or achievement?
- CANAL -- Cyber Activity News Alerting Language Model: Empirical Approach vs. Expensive LLM
- Who's in and who's out? A case study of multimodal CLIP-filtering in DataComp
- Democratizing LLMs: An Exploration of Cost-Performance Trade-offs in Self-Refined Open-Source Models
- GPT Struct Me: Probing GPT Models on Narrative Entity Extraction
- CNNs Avoid Curse of Dimensionality by Learning on Patches
- Complex QA and language models hybrid architectures, Survey
- Generalist Models in Medical Image Segmentation: A Survey and Performance Comparison with Task-Specific Approaches
- PartIR: Composing SPMD Partitioning Strategies for Machine Learning
- Reward-Augmented Decoding: Efficient Controlled Text Generation With a Unidirectional Reward Model
- On the average-case complexity of learning output distributions of quantum circuits
- Variational Open-Domain Question Answering
- Neurosymbolic Information Extraction from Transactional Documents
- Efficient Training of Self-Supervised Speech Foundation Models on a Compute Budget
- The LLM Mirage: Economic Interests and the Subversion of Weaponization Controls
- Large Language Models -- the Future of Fundamental Physics?
- Simple and Effective Input Reformulations for Translation
- Hallucination Detection in Large Language Models Using Diversion Decoding