3 papers
cs.LG2026
Grouter: Decoupling Routing from Representation for Accelerated MoE Training
Yuqi Xu, Rizhen Hu, Zihan Liu +2
Traditional Mixture-of-Experts (MoE) training typically proceeds without any structural priors, effectively requiring the model to simultaneously train expert weights while searchi…
math.OC2025
Clapping: Removing Per-sample Storage for Pipeline Parallel Distributed Optimization with Communication Compression
Boao Kong, Xu Huang, Yuqi Xu +3
Pipeline-parallel distributed optimization is essential for large-scale machine learning but is challenged by significant communication overhead from transmitting high-dimensional…
cs.LG2025
An Efficient Subspace Algorithm for Federated Learning on Heterogeneous Data
Jiaojiao Zhang, Yuqi Xu, Kun Yuan
This work addresses the key challenges of applying federated learning to large-scale deep neural networks, particularly the issue of client drift due to data heterogeneity across c…