activity
20242026
most citedScaling up Masked Diffusion Models on Text

1 citations · 1 across the 6 of their papers we have counts for

collaborators

8 papers

cs.AI2026

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models

Fengqi Zhu, Shaoxuan Xu, Jingyang Ou +11

Diffusion language models (dLLMs) offer an alternative to autoregressive (AR) language modeling, yet the scaling behavior of Mixture-of-Experts (MoE) dLLMs remains poorly understoo…

cs.CV2026

AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution

Zehao Wang, Yihan Zeng, Zidong Gong +5

Post-training via Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) is crucial for enhancing reasoning in Multimodal Large Language Models (MLLMs), yet existing paradigm…

cs.CL2025

UltraLLaDA: Scaling the Context Length to 128K for Diffusion Large Language Models

Guangxin He, Shen Nie, Fengqi Zhu +6

Diffusion LLMs have attracted growing interest, with plenty of recent work emphasizing their great potential in various downstream tasks; yet the long-context behavior of diffusion…

cs.CL2025

LLaDA-MoE: A Sparse MoE Diffusion Language Model

Fengqi Zhu, Zebin You, Yipeng Xing +23

We introduce LLaDA-MoE, a large language diffusion model with the Mixture-of-Experts (MoE) architecture, trained from scratch on approximately 20T tokens. LLaDA-MoE achieves compet…

cs.LG2025

LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion Models

Fengqi Zhu, Rongzhen Wang, Shen Nie +8

While Masked Diffusion Models (MDMs), such as LLaDA, present a promising paradigm for language modeling, there has been relatively little effort in aligning these models with human…

cs.CL2025

Large Language Diffusion Models

Shen Nie, Fengqi Zhu, Zebin You +7

The capabilities of large language models (LLMs) are widely regarded as relying on autoregressive models (ARMs). We challenge this notion by introducing LLaDA, a diffusion model tr…