Showing cs.DCShow all
2 papers · 1 filter
cs.DC2025
GPU-Initiated Networking for NCCL
Khaled Hamidouche, John Bachan, Pak Markthub +6
Modern AI workloads, especially Mixture-of-Experts (MoE) architectures, increasingly demand low-latency, fine-grained GPU-to-GPU communication with device-side control. Traditional…
cs.DC2024
Optimizing Distributed ML Communication with Fused Computation-Collective Operations
Kishore Punniyamurthy, Khaled Hamidouche, Bradford M. Beckmann
In order to satisfy their ever increasing capacity and compute requirements, machine learning models are distributed across multiple nodes using numerous parallelism strategies. As…