3 papers
cs.LG2026
Olmo Hybrid: From Theory to Practice and Back
William Merrill, Yanhong Li, Tyler Romero +19
Recent work has demonstrated the potential of non-transformer language models, especially linear recurrent neural networks (RNNs) and hybrid models that mix recurrence and attentio…
cs.LG2026
A Theory of Training Profit-Optimal LLMs
Sophie Hao, William Merrill
Scaling LLMs requires tremendous computational resources, and recent advances in AI have gone hand in hand with massive amounts of capital expenditure. While it is established that…
cs.LG2025
Critical Batch Size Revisited: A Simple Empirical Approach to Large-Batch Language Model Training
William Merrill, Shane Arora, Dirk Groeneveld +1
The right batch size is important when training language models at scale: a large batch size is necessary for fast training, but a batch size that is too large will harm token effi…