2 papers
cs.LG2026
Broken Symmetry in BF16 Attention: Why FlashAttention Gradients Blow Up Late in Training
Junlin Chen, Daize Dong, Huanwei Di +9
BF16 is now standard in large-scale pretraining, including in fused attention kernels such as FlashAttention, and these kernels are widely trusted. When we used FlashAttention-3 to…
cs.LG2026
PR2: Predictive Routing Replay for MoE-Based LLM Reinforcement Learning
Daize Dong, Junlin Chen, Haolong Jia +9
Mixture of Experts (MoE) Large Language Models (LLMs) achieve strong performance at scale. However, reinforcement learning (RL) on MoE-based LLMs often suffers from training instab…