3 papers
cs.LG2025
Critical Batch Size Revisited: A Simple Empirical Approach to Large-Batch Language Model Training
William Merrill, Shane Arora, Dirk Groeneveld +1
The right batch size is important when training language models at scale: a large batch size is necessary for fast training, but a batch size that is too large will harm token effi…
cs.CL2025
2 OLMo 2 Furious
Team OLMo, Pete Walsh, Luca Soldaini +40
We present OLMo 2, the next generation of our fully open language models. OLMo 2 includes a family of dense autoregressive language models at 7B, 13B and 32B scales with fully rele…
math.FA2020
Solution branches of nonlinear eigenvalue problems on restricted domains
Shane Arora
We extend bifurcation results of nonlinear eigenvalue problems from real Banach spaces to any neighbourhood of a given point. For points of odd multiplicity on these restricted dom…