1 paper
Bhavesh Kumar, Roger Jin, Jeffrey Quesnelle
As language models scale to trillions of parameters, distributed training across many GPUs becomes essential, yet gradient synchronization over high-bandwidth, low-latency networks…