5 papers
Cross-Tokenizer On-Policy Distillation via Byte-Prefix Marginalization
Hao Wang, Kun Yuan, Wenlin Zhong +4
Open-weight language models from different families exhibit complementary capabilities, motivating their consolidation into a compact student through on-policy distillation (OPD).…
GNMR: Runtime Stability Control for Low-Precision Large Language Model Training
Boao Kong, Weichen Jia, Engao Zhang +6
Training stability is a key bottleneck in low-precision language model training: efficient low-cost paths can still produce short-lived numerical risks at a small set of operators.…
Row-Stochastic Matrices Can Provably Outperform Doubly Stochastic Matrices in Decentralized Learning
Bing Liu, Boao Kong, Limin Lu +2
Decentralized learning often involves a weighted global loss with heterogeneous node weights . We revisit two natural strategies for incorporating these weights: (i) embedding t…
BROS: Bias-Corrected Randomized Subspaces for Memory-Efficient Single-Loop Bilevel Optimization
Hengrui Zhang, Boao Kong, Engao Zhang +1
Stochastic bilevel optimization (SBO) has become a standard framework for hyperparameter learning, data reweighting, representation learning, and data-mixture optimization in deep…
Synergistic Intra- and Cross-Layer Regularization Losses for MoE Expert Specialization
Rizhen Hu, Yuan Cao, Boao Kong +2
Sparse Mixture-of-Experts (MoE) models scale Transformers efficiently but suffer from expert overlap -- redundant representations across experts and routing ambiguity, resulting in…