credit assignment 2agentic reinforcement learning 1agentic RL 1context management 1large language models 1long-horizon tasks 1memory compression 1self-distillation 1verifiable rewards 1
From the 2 of 6 linked papers with an AI index.
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Group-Reflective Self-Distillation for Agentic Reinforcement Learning
Binbin Zheng, Zijun Xie, Guanqun Zhao +4
The paper introduces Group-Reflective Self-Distillation (GRSD), a method that uses a policy's own verified rollouts to generate privileged guidance for better credit assignment in…
cs.AI2026
Deconstructing Off-Policy Ratios: Entropy-Scaled Trust Regions for Asynchronous Reinforcement Learning
Guanqun Zhao, Zijun Xie, Binbin Zheng +5
Asynchronous reinforcement learning (RL) accelerates large language model (LLM) post-training by overlapping rollout generation with policy optimization, but the resulting stale, o…