activity
20242026
collaborators

12 papers

cs.LG2026

Hyperball May Not Be a Free Lunch

Yihao Xiao, Jialong Sun, Zitian Gao +5

For scale-invariant deep networks, Hyperball-style optimizers have shown strong performance in large-scale training by fixing the norms of matrix-valued parameters and normalizing…

cs.CL2026

Loop the Loopies!

Zitian Gao, Yilong Chen, Yihao Xiao +4

We present the Loopie series, consisting of two Mixture-of-Experts (MoE) models: a 20B-parameter model with 2B active parameters and a 6B-parameter model with 0.6B active parameter…

cs.CL2026

Beyond N-gram: Data-Aware X-GRAM Extraction for Efficient Embedding Parameter Scaling

Yilong Chen, Yanxi Xie, Zitian Gao +10

Large token-indexed lookup tables provide a compute-decoupled scaling path, but their practical gains are often limited by poor parameter efficiency and rapid memory growth. We att…

cs.AI2025

Universal Reasoning Model

Zitian Gao, Lynx Chen, Yihao Xiao +5

Universal transformers (UTs) have been widely used for complex reasoning tasks such as ARC-AGI and Sudoku, yet the specific sources of their performance gains remain underexplored.…

cs.CL2025

What Makes Diffusion Language Models Super Data Learners?

Zitian Gao, Haoming Luo, Lynx Chen +4

Recent studies have shown that diffusion language models achieve remarkable data efficiency under limited-data constraints, yet the underlying mechanisms remain unclear. In this wo…

cs.CL2025

One-shot Entropy Minimization

Zitian Gao, Lynx Chen, Haoming Luo +2

We trained 13,440 large language models and found that entropy minimization requires only a single unlabeled data and 10 steps optimization to achieve performance improvements comp…