1 paper
Ertza Warraich, Ali Imran, Annus Zulfiqar +3
As distributed machine learning (ML) workloads scale to thousands of GPUs connected by high-speed interconnects, tail latency in collective communication has become a major bottlen…