3 papers
cs.CV2026
ExFusion: Efficient Transformer Training via Multi-Experts Fusion
Jiacheng Ruan, Daize Dong, Xiaoye Qu +5
Mixture-of-Experts (MoE) models substantially improve performance by increasing the capacity of dense architectures. However, directly training MoE models requires considerable com…
cs.LG2026
PR2: Predictive Routing Replay for MoE-Based LLM Reinforcement Learning
Daize Dong, Junlin Chen, Haolong Jia +9
Mixture of Experts (MoE) Large Language Models (LLMs) achieve strong performance at scale. However, reinforcement learning (RL) on MoE-based LLMs often suffers from training instab…
cs.LG2025
Towards Efficient Mixture of Experts: A Holistic Study of Compression Techniques
Shwai He, Daize Dong, Liang Ding +1
Scaling large language models has driven remarkable advancements across various domains, yet the continual increase in model size presents significant challenges for real-world dep…