1 paper
David McAllister, Matthew Tancik, Jiaming Song +1
Large-scale AI model training divides work across thousands of GPUs, then synchronizes gradients across them at each step. This incurs a significant network burden that only centra…