2 papers
cs.LG2026
Efficient RLVR Scheduling via Graph-Structured Online Difficulty Estimation
Zhizhao Liu, Zhiliang Tian, Xi Wang +4
Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models but relies on costly rollout exploration. Assigning the same expl…
cs.CL2026
A Single Rewrite Suffices: Empirical Lessons from Production Skill Description Optimization
Yangqiaoyu Zhou, Mohammad Alqudah, Kwei-Herng Lai +3
Enterprise AI agents route user queries to specialized skills by matching queries against natural language skill descriptions. When two skills share overlapping descriptions, the r…