works on

From the 3 of 11 linked papers with an AI index.

collaborators

11 papers

cs.CL2026

TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning

Changle Qu, Sunhao Dai, Hengyi Cai +4

Tool-Integrated Reasoning (TIR) enables LLMs to solve complex tasks through iterative tool interactions. However, existing reinforcement learning methods often rely on trajectory-l…

cs.LG2026

SCOPE-RL: Optimizing Reasoning Paths Before and After Success

Xiaojian Liu, Han Xu, Jianqiang Xia +6

The paper proposes SCOPE-RL, a two-stage reinforcement learning framework that adds dense, verifiable rewards to both pre‑success and post‑success reasoning steps of large language…

cs.LG2026

Tracking Drift: Variation-Aware Entropy Scheduling for Non-Stationary Reinforcement Learning

Tongxi Wang, Zhuoyang Xia, Xinran Chen +1

The paper proposes an adaptive method for adjusting the entropy coefficient in reinforcement learning to handle non‑stationary environments, using online drift proxies to scale exp…

cs.AI2026

STAMP: Provenance-Guided Credit Assignment for Deep Search Agents

Ke Xu, Han Xu, Xinran Chen +6

The paper presents STAMP, a method that assigns credit to individual actions of deep search agents by verifying whether retrieved documents support evidence in a training-time grap…

cs.CL2026

Measuring Maximum Activations in Open Large Language Models

Luxuan Chen, Han Tian, Xinran Chen +9

The dynamic range of activations is a first-order constraint for low-bit quantization, activation scaling, and stable LLM inference. Prior work characterized outlier features and m…

cs.CL2026

EndPrompt: Efficient Long-Context Extension via Terminal Anchoring

Han Tian, Luxuan Chen, Xinran Chen +10

Extending the context window of large language models typically requires training on sequences at the target length, incurring quadratic memory and computational costs that make lo…