2 papers
cs.DC2026
SpCCL: A Sparsity-Aware Collective Communication Library for GPU Platforms
Lannie Dalton Hough, Emir Gencer, Hoffmann Muki +1
Collective communication is essential to high performance computing and machine learning workloads, yet libraries such as NCCL do not exploit sparsity in message payloads. Sending…
cs.DC2025
Understanding and Improving Communication Performance in Multi-node LLM Inference
Prajwal Singhania, Siddharth Singh, Lannie Dalton Hough +4
As large language models (LLMs) continue to grow in size, distributed inference has become increasingly important. Model-parallel strategies must now efficiently scale not only acr…