2 citations · 2 across the 1 of their papers we have counts for
1 paper · 1 filter
William Merrill, Shane Arora, Dirk Groeneveld +1
The right batch size is important when training language models at scale: a large batch size is necessary for fast training, but a batch size that is too large will harm token effi…