2 papers
cs.CL2026
Data-Efficient Autoregressive-to-Diffusion Language Models via On-Policy Distillation
Xingyu Su, Jacob Helwig, Shubham Parashar +6
We study the transformation of autoregressive models (ARLMs) into diffusion language models (DLMs). Rather than pretraining from scratch, prior work replaces the causal attention i…
cs.CL2026
Learnability-Informed Fine-Tuning of Diffusion Language Models
Shubham Parashar, Atharv Chagi, Jacob Helwig +5
We aim to improve the reasoning capabilities of diffusion language models (DLMs). While SFT is a popular post-training recipe for autoregressive models, its use in DLMs faces chall…