Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Fantastic Pretraining Optimizers and Where to Find Them II: Hyperball Optimization
Kaiyue Wen, Xingyu Dang, Kaifeng Lyu +2
Matrix based optimizers such as Muon can substantially speed up language model pretraining, but their gains over AdamW are observed to shrink as model size and data scale grow when…
cs.LG2026
Escaping the Cognitive Well: Efficient Competition Math with Off-the-Shelf Models
Xingyu Dang, Rohit Agarwal, Rodrigo Porto +3
In the past year, custom and unreleased math reasoning models reached gold medal performance on the International Mathematical Olympiad (IMO). Similar performance was then reported…
cs.LG2025
Weight Ensembling Improves Reasoning in Language Models
Xingyu Dang, Christina Baek, Kaiyue Wen +2
We investigate a failure mode that arises during the training of reasoning models, where the diversity of generations begins to collapse, leading to suboptimal test-time scaling. N…