3 papers
cs.LG2026
Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization
Yixuan Wang, Yifei Chen, Haichao Zhang +6
Reinforcement learning (RL) with group-relative advantages has become the de facto standard for post-training language model reasoners. However, when optimizing multiple reward obj…
cs.AI2026
Demystifying Agent Skills: Why They Work-Until They Don't
Zhiyuan Jiang, Fangrui Huang, Hanwen Xing +6
Skills have emerged as a practical and effective approach for enhancing LLM agents at inference time through structured packages of knowledge. However, existing evaluations largely…
cs.DC2025
OSGym: Scalable OS Infra for Computer Use Agents
Zengyi Qin, Jinyuan Chen, Yunze Man +25
Training computer use agents requires full-featured OS sandboxes with GUI environments, which consume substantial hardware resources as the number of sandboxes scales. Stochastic e…