From the 2 of 5 linked papers with an AI index.
5 papers
Agentic Reinforcement Learning with Self-Distilled Reward Shaping
Ranxu Zhang, Guinan Chen, Chenshaodong +5
Agentic reinforcement learning enables LLM agents to learn through interaction, but sparse trajectory-level rewards reveal success without identifying which intermediate decisions…
When Does Muon Help Agentic Reinforcement Learning?
Kai Ruan, Jinghao Lin, Zihe Huang +4
Muon is competitive with AdamW in large-scale pre-training, but its operating regime in reinforcement-learning post-training remains unclear. We map this regime on ALFWorld, a spar…
Doomed from the Start: Early Abort of LLM Agent Episodes via a Recall-Controlled Probe Cascade
Kai Ruan, Zihe Huang, Ziqi Zhou +4
The paper proposes using lightweight linear probes on hidden states of large language model agents to predict failures early and abort doomed episodes, achieving large compute savi…
The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy
Chunzheng Zhu, Lei Tian, Bohan Tan +16
The paper outlines a roadmap for developing medical AI agents that move from assisting clinicians to operating autonomously, focusing on scaling frameworks, capabilities, and clini…
GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy
Hongze Tan, Zihan Wang, Jianfei Pan +7
Reinforcement Learning (RL) is pivotal for enhancing Large Language Model (LLM) reasoning, yet mainstream algorithms such as GRPO and DAPO remain constrained by a coarse-grained cr…