2 papers
cs.AI2026
ARISE-RL: Agentic Rubric-Grounded Iterative Self-Evolution with Reinforcement Learning
Fanrui Zhang, Ruixue Ding, Qiang Zhang +13
Training open-ended agents via reinforcement learning (RL) is hindered by the lack of verifiable gold answers and scalable rubrics. Moreover, even near the model's capability bound…
cs.LG2026
MemoryWalker: Stop Training Agents on Contexts They Never Saw
Zinco J, Xunjie Zhu, Shen Huang +3
Production agent harnesses such as Claude Code and Qwen-Agent compress context during rollout, but training under compression creates a conditioning problem: every eviction branche…