2 papers
cs.AI2026
EvoCUA-1.5: Online Reinforcement Learning for Multi-turn Computer-Use Agents
Mianqiu Huang, Taofeng Xue, Chong Peng +12
Computer-use agents must solve long-horizon tasks through repeated interaction with partially observable, multimodal desktop environments. Although imitation learning and offline t…
cs.CL2026
Retrospective Progress-Aware Self-Refinement for LLM Agent Training
Xinbei Ma, Congmin Zheng, Jiyang Qiu +10
LLM-based agents trained with reinforcement learning optimize step-wise action prediction but lack metacognitive awareness of task progress, inducing a gap that hinders long-horizo…