Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
SAF-OPD: Stable Advantage Fusion for On-Policy Distillation
Yifan Ding, Xincheng Wei, Yoshua Y. Li +7
Reinforcement learning with verifiable rewards (RLVR) broadcasts a single response-level reward to every token, while on-policy distillation (OPD) scores each token against a stron…
cs.LG2026
SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution
Zhiyuan Yao, Yuxin Chen, Zhengxi Lu +13
Large language model agents often encounter related yet distinct tasks that share reusable solution patterns. Yet standard agentic reinforcement learning treats tasks as independen…
cs.LG2026
EDIS: Diagnosing LLM Reasoning via Entropy Dynamics
Chenghua Zhu, Siyan Wu, Xiangkang Zeng +6
Entropy-based confidence signals are increasingly leveraged to improve reasoning in large language models (LLMs), yet existing approaches treat confidence as a static quantity -- t…