#on-policy distillation

topicon-policy distillation

14 papers · 1 filter

cs.CL2026

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models

Yecheng Wu, Song Han, Han Cai

The paper proposes Lightning OPD 2.0, a method that reduces style‑related bias when using on‑policy distillation across different teacher models, improving performance on mathemati…

cs.AI2026

Correcting What You Cannot See: Credit Assignment for Perception Distillation in Multimodal Reasoners

Feng Xiong, Leyan Xue, Hongyu Lin

The paper proposes Perception-Correction Distillation (PCD), a label‑free method that uses downstream failures and teacher‑student disagreement to pinpoint and correct perception e…

cs.LG2026

Flux-OPD: On-Policy Distillation with Evolving Contexts

Yuran Wang, Zekun Wang, Bohan Zeng +10

The paper introduces Flux-OPD, a method for training large language models by distilling knowledge from teachers while using evolving contexts as supervision, and stabilizes the pr…

cs.CV2026

Veritas++: Value-aware On-Policy Distillation for Perception-Enhanced AIGI Detection

Hao Tan, Jun Lan, Zichang Tan +7

Veritas++ introduces a perception‑enhanced framework for detecting AI‑generated images by training models to capture fine‑grained visual details, semantic anomalies, and pixel‑leve…

cs.AI2026

On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment

Yongjian Guo, Wanlun Ma, Lingyu Shen +2

The paper introduces Routing-based On-Policy Distillation (ROPD), a method for safely realigning large language models that resists malicious prompt templates while preserving the…

cs.LG2026

Weak-to-Strong On-Policy Distillation

Fangxu Yu, Zinan Lin, Xiaodong Liu +4

The paper proposes Weak-to-Strong On-Policy Distillation (W2S-OPD), a method that improves a large language model by distilling knowledge from multiple weaker models using a constr…