3 papers
cs.CL2026
Roll Out and Roll Back: Diffusion LLMs are Their Own Efficiency Teachers
Fanqin Zeng, Feng Hong, Geng Yu +6
Diffusion Large Language Models (DLLMs) promise fast parallel generation, yet open-source DLLMs still face a severe quality-speed trade-off: accelerating decoding by revealing mult…
cs.LG2025
Innovator: Scientific Continued Pretraining with Fine-grained MoE Upcycling
Ning Liao, Xiaoxing Wang, Zehao Lin +18
A large language model (LLM) with knowledge in both scientific and general tasks is the foundation of science general intelligence. However, directly continued pretraining an LLM u…
cs.CL2025
Wide-In, Narrow-Out: Revokable Decoding for Efficient and Effective DLLMs
Feng Hong, Geng Yu, Yushi Ye +5
Diffusion Large Language Models (DLLMs) have emerged as a compelling alternative to Autoregressive models, designed for fast parallel generation. However, existing DLLMs are plague…