Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
SPO++: Stream-Aligned Policy Optimization for Asynchronous Agentic RL
Kai Ruan, Jinghao Lin, Qianshan Wei +2
Group-relative reinforcement learning waits for sibling rollouts of the same prompt, which is costly for long and variable tool-use trajectories. Single-stream Policy Optimization…
cs.AI2026
Doomed from the Start: Early Abort of LLM Agent Episodes via a Recall-Controlled Probe Cascade
Kai Ruan, Zihe Huang, Ziqi Zhou +4
Large language model (LLM) agents often waste inference compute by continuing multi-step trajectories that are already doomed to fail. We study early failure prediction and inferen…