Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe
Wenjin Hou, Shangpin Peng, Weinong Wang +13
On-policy distillation (OPD) has recently emerged as an effective post-training paradigm for consolidating the capabilities of specialized expert models into a single student model…
cs.LG2025
Agentic Reinforced Policy Optimization
Guanting Dong, Hangyu Mao, Kai Ma +11
Large-scale reinforcement learning with verifiable rewards (RLVR) has demonstrated its effectiveness in harnessing the potential of large language models (LLMs) for single-turn rea…