2 papers
cs.CL2026
TRACE: A Self-Evolving Skill Bank for Consistent, Limit-Aware LLM Agents
Wenhao Wu, Menghao Zhang, Xin Wang +3
Reliable deployment of LLM agents in user-facing products depends not on raw task-solving ability but on consistency and limit-awareness: behaving the same way across repeated tria…
cs.DC2026
PlexRL: Cluster-Level Orchestration of Serviceized LLM Execution for RLVR
Yiqi Zhang, Fangzheng Jiao, Tian Tang +13
Reinforcement learning with verifiable rewards (RLVR) has recently unlocked strong reasoning capabilities in large language models (LLMs), triggering rapid exploration of new algor…