Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Self-Supervised Skill Optimization
Siran Peng, Cuiyu Yang, Tianyu Fu +9
Agent skills provide frozen large language model (LLM) agents with reusable procedural guidance, and recent work shows that such skills can be optimized with ground-truth (GT) feed…
cs.CL2026
One Ring to Rule Them All: Unifying Group-Based RL via Dynamic Power-Mean Geometry
Weisong Zhao, Tong Wang, Zichang Tan +11
Group-based reinforcement learning has evolved from the arithmetic mean of GRPO to the geometric mean of GMPO. While GMPO improves stability by constraining a conservative objectiv…