4 papers
MemFine: Memory-Aware Fine-Grained Scheduling for MoE Training
Lu Zhao, Rong Shi, Shaoqing Zhang +21
The training of large-scale Mixture of Experts (MoE) models faces a critical memory bottleneck due to severe load imbalance caused by dynamic token routing. This imbalance leads to…
Every Activation Boosted: Scaling General Reasoner to 1 Trillion Open Language Foundation
Ling Team, Ang Li, Ben Liu +138
We introduce Ling 2.0, a series reasoning-oriented language foundation built upon the principle that every activation boosts reasoning capability. Designed to scale from tens of bi…
Plan before Solving: Problem-Aware Strategy Routing for Mathematical Reasoning with LLMs
Shihao Qi, Jie Ma, Ziang Yin +5
Existing methods usually leverage a fixed strategy, such as natural language reasoning, code-augmented reasoning, tool-integrated reasoning, or ensemble-based reasoning, to guide L…
From Static to Dynamic: Adaptive Monte Carlo Search for Mathematical Process Supervision
Jie Ma, Shihao Qi, Rui Xing +4
The quality of process data plays a key role in training a Process Reward Model (PRM), which can enhance the complex mathematical reasoning capability of large language models. Exi…