Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Training-Trajectory-Aware Token Selection
Zhanming Shen, Jiaqi Hu, Zeyu Qin +7
Efficient distillation is a key pathway for converting expensive reasoning capability into deployable efficiency, yet in the frontier regime where the student already has strong re…
cs.CL2026
Parallelism and Generation Order in Masked Diffusion Language Models: Limits Today, Potential Tomorrow
Yangyang Zhong, Yanmei Gu, Zhengqing Zang +14
Masked Diffusion Language Models (MDLMs) promise parallel token generation and arbitrary-order decoding, yet it remains unclear to what extent current models truly realize these ca…