2 papers
cs.DC2026
NEURON-Fabric: Architecture-Runtime Co-Design for Controlled Low-Bit Gradient Communication
Ziqiang Wang, Changcheng Huang, Chung-Horng Lung
Large-scale neural-network training repeatedly aggregates gradients across devices, making communication a central cost in distributed learning. Low-bit gradient aggregation can re…
cs.DC2026
NEURON-Fabric: CXL-Side Low-Bit Gradient Aggregation for Distributed Training
Ziqiang Wang, Changcheng Huang, Chung-Horng Lung
In large-model distributed training, especially large language model workloads, gradient All-Reduce increasingly stresses the memory and communication path. This paper asks whether…