From the 2 of 4 linked papers with an AI index.
4 papers
SCOPE-RL: Optimizing Reasoning Paths Before and After Success
Xiaojian Liu, Han Xu, Jianqiang Xia +6
The paper proposes SCOPE-RL, a two-stage reinforcement learning framework that adds dense, verifiable rewards to both pre‑success and post‑success reasoning steps of large language…
STAMP: Provenance-Guided Credit Assignment for Deep Search Agents
Ke Xu, Han Xu, Xinran Chen +6
The paper presents STAMP, a method that assigns credit to individual actions of deep search agents by verifying whether retrieved documents support evidence in a training-time grap…
Let the Data Decide: Supervision Analysis, Capability Trade-offs, and Adaptive Objective Routing in Continued Pre-Training via Off-Policy Distillation
Jiangan Yuan, Zhixuan Li, Han Xu
Off-policy distillation is now central to large language model pre-training, yet how training data, objective parameterization, and model capabilities interact remains poorly chara…
Data and Evaluation Closed-Loop for Model Capability Enhancement
Zhixuan Li, Jiangan Yuan, Han Xu
Model capability is the central variable in LLM pre-training, yet is never observed directly: data shapes it prospectively, while evaluation reveals it only retrospectively, compre…