4 citations · 4 across the 1 of their papers we have counts for
1 paper
Sami Alabed, Daniel Belov, Bart Chrzaszcz +14
Training of modern large neural networks (NN) requires a combination of parallelization strategies encompassing data, model, or optimizer sharding. When strategies increase in comp…