3 papers
cs.LG2026
Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts Inference
Baihui Liu, Kaiyuan Tian, Wei Wang +3
Mixture-of-Experts (MoE) has become a dominant architecture for scaling large language models due to their sparse activation mechanism. However, the substantial number of expert ac…
cs.AI2026
State-Action Inpainting Diffuser for Continuous Control with Delay
Dongqi Han, Wei Wang, Enze Zhang +1
Signal delay poses a fundamental challenge in continuous control and reinforcement learning (RL) by introducing a temporal gap between interaction and perception. Current solutions…
cs.DC2024
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models
Wei Wang, Zhiquan Lai, Shengwei Li +5
The size of deep learning models has been increasing to enhance model quality. The linear increase in training computation budget with model size means that training an extremely l…