3 papers
cs.LG2026
Fantastic Pretraining Optimizers and Where to Find Them II: Hyperball Optimization
Kaiyue Wen, Xingyu Dang, Kaifeng Lyu +2
Matrix based optimizers such as Muon can substantially speed up language model pretraining, but their gains over AdamW are observed to shrink as model size and data scale grow when…
cs.LG2025
Weight Ensembling Improves Reasoning in Language Models
Xingyu Dang, Christina Baek, Kaiyue Wen +2
We investigate a failure mode that arises during the training of reasoning models, where the diversity of generations begins to collapse, leading to suboptimal test-time scaling. N…
cs.CL2025
Overtrained Language Models Are Harder to Fine-Tune
Jacob Mitchell Springer, Sachin Goyal, Kaiyue Wen +5
Large language models are pre-trained on ever-growing token budgets under the assumption that better pre-training performance translates to improved downstream models. In this work…