2 citations · 2 across the 2 of their papers we have counts for
1 paper · 1 filter
Ertza Warraich, Ali Imran, Annus Zulfiqar +3
As distributed machine learning (ML) workloads scale to thousands of GPUs connected by ultra-high-speed inter-connects, tail latency in collective communication has emerged as a pr…