From the 3 of 11 linked papers with an AI index.
11 papers
TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning
Changle Qu, Sunhao Dai, Hengyi Cai +4
Tool-Integrated Reasoning (TIR) enables LLMs to solve complex tasks through iterative tool interactions. However, existing reinforcement learning methods often rely on trajectory-l…
SCOPE-RL: Optimizing Reasoning Paths Before and After Success
Xiaojian Liu, Han Xu, Jianqiang Xia +6
The paper proposes SCOPE-RL, a two-stage reinforcement learning framework that adds dense, verifiable rewards to both pre‑success and post‑success reasoning steps of large language…
Tracking Drift: Variation-Aware Entropy Scheduling for Non-Stationary Reinforcement Learning
Tongxi Wang, Zhuoyang Xia, Xinran Chen +1
The paper proposes an adaptive method for adjusting the entropy coefficient in reinforcement learning to handle non‑stationary environments, using online drift proxies to scale exp…
STAMP: Provenance-Guided Credit Assignment for Deep Search Agents
Ke Xu, Han Xu, Xinran Chen +6
The paper presents STAMP, a method that assigns credit to individual actions of deep search agents by verifying whether retrieved documents support evidence in a training-time grap…
Measuring Maximum Activations in Open Large Language Models
Luxuan Chen, Han Tian, Xinran Chen +9
The dynamic range of activations is a first-order constraint for low-bit quantization, activation scaling, and stable LLM inference. Prior work characterized outlier features and m…
EndPrompt: Efficient Long-Context Extension via Terminal Anchoring
Han Tian, Luxuan Chen, Xinran Chen +10
Extending the context window of large language models typically requires training on sequences at the target length, incurring quadratic memory and computational costs that make lo…