117 citations · 141 across the 4 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2022★ 7 cited
lo-fi: distributed fine-tuning without communication
Mitchell Wortsman, Suchin Gururangan, Shen Li +4
When fine-tuning large neural networks, it is common to use multiple nodes and to communicate gradients at each optimization step. By contrast, we investigate completely local fine…
cs.LG2021★ 16 cited
PipeTransformer: Automated Elastic Pipelining for Distributed Training of Transformers
Chaoyang He, Shen Li, Mahdi Soltanolkotabi +1
The size of Transformer models is growing at an unprecedented pace. It has only taken less than one year to reach trillion-level parameters after the release of GPT-3 (175B). Train…