Showing 2026Show all
2 papers · 1 filter
cs.CV2026
ExFusion: Efficient Transformer Training via Multi-Experts Fusion
Jiacheng Ruan, Daize Dong, Xiaoye Qu +5
Mixture-of-Experts (MoE) models substantially improve performance by increasing the capacity of dense architectures. However, directly training MoE models requires considerable com…
cs.LG2026
PR2: Predictive Routing Replay for MoE-Based LLM Reinforcement Learning
Daize Dong, Junlin Chen, Haolong Jia +9
Mixture of Experts (MoE) Large Language Models (LLMs) achieve strong performance at scale. However, reinforcement learning (RL) on MoE-based LLMs often suffers from training instab…