1 paper
Amel Fatima, Tuan Ta, Bradford M. Beckmann
Distributed ML workloads rely heavily on collective communication across multi-GPU, multi-node systems. Emerging scale-up fabrics, such as NVLink and UALink, enable direct memory a…