7 papers
Learning What to Remember: Test-Time Training via Context Distillation
Zixuan Wang, Xingyu Dang, Rui-Jie Zhu +4
Effective long-context modeling is not merely about retaining more of the past, but about preserving the information that may prove relevant later. Test-time training (TTT) is an a…
The Power of Power Law: Asymmetry Enables Compositional Reasoning
Zixuan Wang, Xingyu Dang, Jason D. Lee +1
Natural language data follows a power-law distribution, with most knowledge and skills appearing at very low frequency. While a common intuition suggests that reweighting or curati…
Fantastic Pretraining Optimizers and Where to Find Them II: Hyperball Optimization
Kaiyue Wen, Xingyu Dang, Kaifeng Lyu +2
Matrix based optimizers such as Muon can substantially speed up language model pretraining, but their gains over AdamW are observed to shrink as model size and data scale grow when…
Escaping the Cognitive Well: Efficient Competition Math with Off-the-Shelf Models
Xingyu Dang, Rohit Agarwal, Rodrigo Porto +3
In the past year, custom and unreleased math reasoning models reached gold medal performance on the International Mathematical Olympiad (IMO). Similar performance was then reported…
Goedel-Architect: Streamlining Formal Theorem Proving with Blueprint Generation and Refinement
Jui-Hui Chung, Ziyang Cai, Zihao Li +14
We introduce Goedel-Architect, an agentic framework for formal theorem proving in Lean 4 centered on blueprint generation and refinement. A blueprint is a dependency graph of defin…
Weight Ensembling Improves Reasoning in Language Models
Xingyu Dang, Christina Baek, Kaiyue Wen +2
We investigate a failure mode that arises during the training of reasoning models, where the diversity of generations begins to collapse, leading to suboptimal test-time scaling. N…