works on

From the 2 of 5 linked papers with an AI index.

collaborators

5 papers

cs.LG2026

Agentic Reinforcement Learning with Self-Distilled Reward Shaping

Ranxu Zhang, Guinan Chen, Chenshaodong +5

Agentic reinforcement learning enables LLM agents to learn through interaction, but sparse trajectory-level rewards reveal success without identifying which intermediate decisions…

cs.LG2026

When Does Muon Help Agentic Reinforcement Learning?

Kai Ruan, Jinghao Lin, Zihe Huang +4

Muon is competitive with AdamW in large-scale pre-training, but its operating regime in reinforcement-learning post-training remains unclear. We map this regime on ALFWorld, a spar…

cs.AI2026

Doomed from the Start: Early Abort of LLM Agent Episodes via a Recall-Controlled Probe Cascade

Kai Ruan, Zihe Huang, Ziqi Zhou +4

The paper proposes using lightweight linear probes on hidden states of large language model agents to predict failures early and abort doomed episodes, achieving large compute savi…

cs.AI2026

The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy

Chunzheng Zhu, Lei Tian, Bohan Tan +16

The paper outlines a roadmap for developing medical AI agents that move from assisting clinicians to operating autonomously, focusing on scaling frameworks, capabilities, and clini…

cs.CL2026

GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy

Hongze Tan, Zihan Wang, Jianfei Pan +7

Reinforcement Learning (RL) is pivotal for enhancing Large Language Model (LLM) reasoning, yet mainstream algorithms such as GRPO and DAPO remain constrained by a coarse-grained cr…