2 papers
cs.DC2026
CCCL: Node-Spanning GPU Collectives with CXL Memory Pooling
Dong Xu, Han Meng, Xinyu Chen +11
Large language models (LLMs) training or inference across multiple nodes introduces significant pressure on GPU memory and interconnect bandwidth. The Compute Express Link (CXL) sh…
cs.NI2024
Design and Optimization of Hierarchical Gradient Coding for Distributed Learning at Edge Devices
Weiheng Tang, Jingyi Li, Lin Chen +1
Edge computing has recently emerged as a promising paradigm to boost the performance of distributed learning by leveraging the distributed resources at edge nodes. Architecturally,…