Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Early-stopping for Transformer model training
Jing He, Hua Jiang, Cheng Li +2
This work, based on Random Matrix Theory (RMT), introduces a novel early-stopping strategy for Transformer training dynamics. Utilizing the Power Law (PL) fit to tansformer attenti…
cs.LG2025
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
Cheng Li, Jiexiong Liu, Yixuan Chen +1
Transformer models based on the Mixture of Experts (MoE) architecture have made significant progress in long-sequence modeling, but existing models still have shortcomings in compu…