works on

From the 1 of 7 linked papers with an AI index.

collaborators

7 papers

cs.AI2026

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models

ZhiYan Hou, Xinyu Tang, Hongyan An +9

Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models using automatically verifiable outcome signals, but these signals…

cs.CL2026

EasyOPD: An Easy-to-use On-Policy Distillation Framework for Large Language Models

Jie Sun, Mao Zheng, Mingyang Song +7

The paper introduces EasyOPD, a modular framework that simplifies on-policy distillation for large language models by separating configuration, supervision logic, and distributed e…

cs.LG2026

On-Policy Distillation with Curriculum Turn-level Guidance for Multi-turn Agents

Gengsheng Li, Mao Zheng, Mingyang Song +8

Multi-turn agents that plan, invoke tools, and interact with environments offer a promising paradigm for solving complex tasks, yet their capabilities typically rely on very large…

cs.CV2026

Visual-Advantage On-Policy Distillation for Vision-Language Models

Ruiqi Liu, Xiaolei Lv, Gengsheng Li +8

On-policy knowledge distillation has proven effective for language models, yet its application to vision-language models (VLMs) remains underexplored. We observe that standard on-p…

cs.LG2026

Rubric-based On-policy Distillation

Junfeng Fang, Zhepei Hong, Mao Zheng +7

On-policy distillation (OPD) is a powerful paradigm for model alignment, yet its reliance on teacher logits restricts its application to white-box scenarios. We contend that struct…

cs.LG2026

Unifying Group-Relative and Self-Distillation Policy Optimization via Sample Routing

Gengsheng Li, Tianyu Yang, Junfeng Fang +6

Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for post-training large language models. While Group Relative Policy Optimization (GRPO) is wid…