1 paper
Wenchen Han, Shay Vargaftik, Michael Mitzenmacher +1
Multi-hop all-reduce is the de facto backbone of large model training. As the training scale increases, the network often becomes a bottleneck, motivating the reduction of the volu…