activity
20242026
collaborators

7 papers

cs.CL2026

Learning What to Remember: Test-Time Training via Context Distillation

Zixuan Wang, Xingyu Dang, Rui-Jie Zhu +4

Effective long-context modeling is not merely about retaining more of the past, but about preserving the information that may prove relevant later. Test-time training (TTT) is an a…

cs.AI2026

The Power of Power Law: Asymmetry Enables Compositional Reasoning

Zixuan Wang, Xingyu Dang, Jason D. Lee +1

Natural language data follows a power-law distribution, with most knowledge and skills appearing at very low frequency. While a common intuition suggests that reweighting or curati…

cs.LG2026

Fantastic Pretraining Optimizers and Where to Find Them II: Hyperball Optimization

Kaiyue Wen, Xingyu Dang, Kaifeng Lyu +2

Matrix based optimizers such as Muon can substantially speed up language model pretraining, but their gains over AdamW are observed to shrink as model size and data scale grow when…

cs.LG2026

Escaping the Cognitive Well: Efficient Competition Math with Off-the-Shelf Models

Xingyu Dang, Rohit Agarwal, Rodrigo Porto +3

In the past year, custom and unreleased math reasoning models reached gold medal performance on the International Mathematical Olympiad (IMO). Similar performance was then reported…

cs.AI2026

Goedel-Architect: Streamlining Formal Theorem Proving with Blueprint Generation and Refinement

Jui-Hui Chung, Ziyang Cai, Zihao Li +14

We introduce Goedel-Architect, an agentic framework for formal theorem proving in Lean 4 centered on blueprint generation and refinement. A blueprint is a dependency graph of defin…

cs.LG2025

Weight Ensembling Improves Reasoning in Language Models

Xingyu Dang, Christina Baek, Kaiyue Wen +2

We investigate a failure mode that arises during the training of reasoning models, where the diversity of generations begins to collapse, leading to suboptimal test-time scaling. N…