Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Step-Level On-Policy Distillation: Interpolating Between On-Policy Distillation and Supervised Fine-Tuning
Changhui Sun, Lanbo Liu, Hang Lei +11
On-policy distillation (OPD) aligns a student model with a teacher's logit distribution on student-generated trajectories. This approach has achieved strong empirical gains and can…
cs.CL2026
MDLMPE: Distribution Aware Positional Encoding for Masked Diffusion Language Models
Tong Ling, Hang Lei, Feng Xiao +5
Masked diffusion language models (MDLMs) enable parallel generation and bidirectional context modeling, but their positional context differs fundamentally from that of autoregressi…