3 papers
cs.DC2026
StrataCL: Fabric-Native Communication Library for Production Supernodes
Tiancheng Hu, Jin Qin, Yuzheng Wang +14
Modern distributed AI workloads run across hundreds of accelerators, making communication a major bottleneck. Existing communication libraries remain largely buffer-centric because…
cs.DC2026
Relay Buffer Independent Communication over Pooled HBM for Efficient MoE Inference on Ascend
Tianlun Hu, Tiancheng Hu, Shengsheng Litang +8
Mixture-of-Experts (MoE) inference requires large-scale token exchange across devices, making dispatch and combine major bottlenecks in both prefill and decode. Beyond network tran…
cs.DC2025
FlashRecovery: Fast and Low-Cost Recovery from Failures for Large-Scale Training of LLMs
Haijun Zhang, Jinxiang Wang, Zhenhua Yu +20
Large language models (LLMs) have made a profound impact across various fields due to their advanced capabilities. However, training these models at unprecedented scales requires e…