2 papers
cs.LG2026
Explore More, Drift Less: Outcome-Only Reinforcement Learning Can Suffice for Long-Horizon Interactive Agents
Liming Pu, Xiaoxia Li, Yifu Liu +2
Reinforcement learning is a natural way to post-train LLM agents for long-horizon interactive tasks judged only by end-of-task verification, yet a shared belief holds that outcome-…
cs.AI2026
APIFlow-Bench: Measuring Whether Agents Survive Long, Dependent API Workflows
Zelin Wan, Arash Nourian, Xiaoxiao Li +2
Tool-using agents are commonly evaluated by a single bit: whether an end-to-end workflow completed. This metric fails to distinguish failures that matter in production, such as exp…