4 papers · 1 filter
EDGE: Experience-Distillation for Guided Exploration in Agentic Reinforcement Learning
Can Xie, Yuyi Zhou, Wen Yang +5
Reinforcement learning with outcome-based objectives such as GRPO enables LLM-based agents to solve complex, long-horizon tasks, yet the reusable exploration patterns embedded in i…
An Efficient and Precise Training Data Construction Framework for Process-supervised Reward Model in Mathematical Reasoning
Wei Sun, Qianlong Du, Fuwei Cui +1
Enhancing the mathematical reasoning capabilities of Large Language Models (LLMs) is of great scientific and practical significance. Researchers typically employ process-supervised…
ChineseWebText 2.0: Large-Scale High-quality Chinese Web Text with Multi-dimensional and fine-grained information
Wanyue Zhang, Ziyong Li, Wen Yang +5
During the development of large language models (LLMs), pre-training data play a critical role in shaping LLMs' capabilities. In recent years several large-scale and high-quality p…
A Survey on Data Selection for LLM Instruction Tuning
Bolin Zhang, Jiahao Wang, Qianlong Du +3
Instruction tuning is a vital step of training large language models (LLMs), so how to enhance the effect of instruction tuning has received increased attention. Existing works ind…