8 papers
LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models
Fengqi Zhu, Shaoxuan Xu, Jingyang Ou +11
Diffusion language models (dLLMs) offer an alternative to autoregressive (AR) language modeling, yet the scaling behavior of Mixture-of-Experts (MoE) dLLMs remains poorly understoo…
LLaDA-o: An Effective and Length-Adaptive Omni Diffusion Model
Zebin You, Xiaolu Zhang, Jun Zhou +2
We present \textbf{LLaDA-o}, an effective and length-adaptive omni diffusion model for multimodal understanding and generation. LLaDA-o is built on a Mixture of Diffusion (MoD) fra…
Effective and Efficient Masked Image Generation Models
Zebin You, Jingyang Ou, Xiaolu Zhang +3
Although masked image generation models and masked diffusion models are designed with different motivations and objectives, we observe that they can be unified within a single fram…
Adaptive Task Balancing for Visual Instruction Tuning via Inter-Task Contribution and Intra-Task Difficulty
Yanqi Dai, Yong Wang, Zebin You +3
Visual instruction tuning is a key training stage of large multimodal models. However, when learning multiple visual tasks simultaneously, this approach often results in suboptimal…
Large Language Diffusion Models
Shen Nie, Fengqi Zhu, Zebin You +7
The capabilities of large language models (LLMs) are widely regarded as relying on autoregressive models (ARMs). We challenge this notion by introducing LLaDA, a diffusion model tr…
LLaDA-MoE: A Sparse MoE Diffusion Language Model
Fengqi Zhu, Zebin You, Yipeng Xing +23
We introduce LLaDA-MoE, a large language diffusion model with the Mixture-of-Experts (MoE) architecture, trained from scratch on approximately 20T tokens. LLaDA-MoE achieves compet…