18 papers
LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models
Fengqi Zhu, Shaoxuan Xu, Jingyang Ou +11
Diffusion language models (dLLMs) offer an alternative to autoregressive (AR) language modeling, yet the scaling behavior of Mixture-of-Experts (MoE) dLLMs remains poorly understoo…
Improved Large Language Diffusion Models
Shen Nie, Qiyang Min, Shaoxuan Xu +7
Modern large language models are predominantly trained with autoregressive factorization and causal attention. We present \emph{iLLaDA}, an 8B masked diffusion language model train…
Spectral Condition for P under Width-Depth Scaling
Chenyu Zheng, Rongzhen Wang, Xinyu Zhang +1
Generative foundation models are increasingly scaled in both width and depth, posing significant challenges for stable feature learning and reliable hyperparameter (HP) transfer ac…
Masked Diffusion Models as Energy Minimization
Sitong Chen, Shen Nie, Jiacheng Sun +4
We present a systematic theoretical framework that interprets masked diffusion models (MDMs) as solutions to energy minimization problems in discrete optimal transport. Specificall…
Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data
Jingyang Ou, Shen Nie, Kaiwen Xue +4
Discrete diffusion models with absorbing processes have shown promise in language modeling. The key quantities to be estimated are the ratios between the marginal probabilities of…
LLaDA-o: An Effective and Length-Adaptive Omni Diffusion Model
Zebin You, Xiaolu Zhang, Jun Zhou +2
We present \textbf{LLaDA-o}, an effective and length-adaptive omni diffusion model for multimodal understanding and generation. LLaDA-o is built on a Mixture of Diffusion (MoD) fra…