7 papers
PithTrain: A Compact and Agent-Native MoE Training System
Ruihang Lai, Hao Kang, Haozhan Tang +6
Mixture-of-Experts (MoE) has become the dominant architecture for frontier language models. To meet this demand, production frameworks have built optimized MoE training stacks over…
Generating Pretraining Tokens from Organic Data for Data-Bound Scaling
Zichun Yu, Chenyan Xiong
LLM pretraining is shifting from a compute-bound to a data-bound regime, where available human (organic) text falls far short of scaling demands. However, reaching the data-bound r…
Efficient Multi-Agent System Training with Data Influence-Oriented Tree Search
Wentao Shi, Zichun Yu, Fuli Feng +2
Monte Carlo Tree Search (MCTS) based methods provide promising approaches for generating synthetic data to enhance the self-training of Large Language Model (LLM) based multi-agent…
RePro: Training Language Models to Faithfully Recycle the Web for Pretraining
Zichun Yu, Chenyan Xiong
High-quality pretraining data is the fossil fuel of large language models (LLMs), yet its reserves are running low for frontier models. In this paper, we introduce RePro, a novel w…
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models
Hao Kang, Zichun Yu, Chenyan Xiong
Recent large language models such as Gemini-1.5, DeepSeek-V3, and Llama-4 increasingly adopt Mixture-of-Experts (MoE) architectures, which offer strong efficiency-performance trade…
MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models
Zichun Yu, Spandan Das, Chenyan Xiong
Pretraining data selection has the potential to improve language model pretraining efficiency by utilizing higher-quality data from massive web data corpora. Current data selection…