29 citations · 30 across the 4 of their papers we have counts for
1 paper · 1 filter
Martin Marek, Sanae Lotfi, Aditya Somasundaram +2
Conventional wisdom dictates that small batch sizes make language model pretraining and fine-tuning unstable, motivating gradient accumulation, which trades off the number of optim…