reinforcement learning 2chain-of-thought 1credit assignment 1drift adaptation 1entropy scheduling 1evidence retrieval 1exploration 1large language models 1non-stationary rl 1online learning 1provenance 1reasoning 1
From the 3 of 11 linked papers with an AI index.
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
SCOPE-RL: Optimizing Reasoning Paths Before and After Success
Xiaojian Liu, Han Xu, Jianqiang Xia +6
The paper proposes SCOPE-RL, a two-stage reinforcement learning framework that adds dense, verifiable rewards to both pre‑success and post‑success reasoning steps of large language…
cs.LG2026
Tracking Drift: Variation-Aware Entropy Scheduling for Non-Stationary Reinforcement Learning
Tongxi Wang, Zhuoyang Xia, Xinran Chen +1
The paper proposes an adaptive method for adjusting the entropy coefficient in reinforcement learning to handle non‑stationary environments, using online drift proxies to scale exp…