collaborators

18 papers

cs.AI2026

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models

Fengqi Zhu, Shaoxuan Xu, Jingyang Ou +11

Diffusion language models (dLLMs) offer an alternative to autoregressive (AR) language modeling, yet the scaling behavior of Mixture-of-Experts (MoE) dLLMs remains poorly understoo…

cs.CL2026

Improved Large Language Diffusion Models

Shen Nie, Qiyang Min, Shaoxuan Xu +7

Modern large language models are predominantly trained with autoregressive factorization and causal attention. We present \emph{iLLaDA}, an 8B masked diffusion language model train…

cs.LG2026

Spectral Condition for P under Width-Depth Scaling

Chenyu Zheng, Rongzhen Wang, Xinyu Zhang +1

Generative foundation models are increasingly scaled in both width and depth, posing significant challenges for stable feature learning and reliable hyperparameter (HP) transfer ac…

cs.LG2026

Masked Diffusion Models as Energy Minimization

Sitong Chen, Shen Nie, Jiacheng Sun +4

We present a systematic theoretical framework that interprets masked diffusion models (MDMs) as solutions to energy minimization problems in discrete optimal transport. Specificall…

cs.LG2026

Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data

Jingyang Ou, Shen Nie, Kaiwen Xue +4

Discrete diffusion models with absorbing processes have shown promise in language modeling. The key quantities to be estimated are the ratios between the marginal probabilities of…

cs.CV2026

LLaDA-o: An Effective and Length-Adaptive Omni Diffusion Model

Zebin You, Xiaolu Zhang, Jun Zhou +2

We present \textbf{LLaDA-o}, an effective and length-adaptive omni diffusion model for multimodal understanding and generation. LLaDA-o is built on a Mixture of Diffusion (MoD) fra…