#on-policy distillation
14 papers · 1 filter
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models
Yecheng Wu, Song Han, Han Cai
The paper proposes Lightning OPD 2.0, a method that reduces style‑related bias when using on‑policy distillation across different teacher models, improving performance on mathemati…
Correcting What You Cannot See: Credit Assignment for Perception Distillation in Multimodal Reasoners
Feng Xiong, Leyan Xue, Hongyu Lin
The paper proposes Perception-Correction Distillation (PCD), a label‑free method that uses downstream failures and teacher‑student disagreement to pinpoint and correct perception e…
Flux-OPD: On-Policy Distillation with Evolving Contexts
Yuran Wang, Zekun Wang, Bohan Zeng +10
The paper introduces Flux-OPD, a method for training large language models by distilling knowledge from teachers while using evolving contexts as supervision, and stabilizes the pr…
Veritas++: Value-aware On-Policy Distillation for Perception-Enhanced AIGI Detection
Hao Tan, Jun Lan, Zichang Tan +7
Veritas++ introduces a perception‑enhanced framework for detecting AI‑generated images by training models to capture fine‑grained visual details, semantic anomalies, and pixel‑leve…
On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment
Yongjian Guo, Wanlun Ma, Lingyu Shen +2
The paper introduces Routing-based On-Policy Distillation (ROPD), a method for safely realigning large language models that resists malicious prompt templates while preserving the…
Weak-to-Strong On-Policy Distillation
Fangxu Yu, Zinan Lin, Xiaodong Liu +4
The paper proposes Weak-to-Strong On-Policy Distillation (W2S-OPD), a method that improves a large language model by distilling knowledge from multiple weaker models using a constr…