1 citations · 1 across the 6 of their papers we have counts for
8 papers
LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models
Fengqi Zhu, Shaoxuan Xu, Jingyang Ou +11
Diffusion language models (dLLMs) offer an alternative to autoregressive (AR) language modeling, yet the scaling behavior of Mixture-of-Experts (MoE) dLLMs remains poorly understoo…
AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution
Zehao Wang, Yihan Zeng, Zidong Gong +5
Post-training via Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) is crucial for enhancing reasoning in Multimodal Large Language Models (MLLMs), yet existing paradigm…
UltraLLaDA: Scaling the Context Length to 128K for Diffusion Large Language Models
Guangxin He, Shen Nie, Fengqi Zhu +6
Diffusion LLMs have attracted growing interest, with plenty of recent work emphasizing their great potential in various downstream tasks; yet the long-context behavior of diffusion…
LLaDA-MoE: A Sparse MoE Diffusion Language Model
Fengqi Zhu, Zebin You, Yipeng Xing +23
We introduce LLaDA-MoE, a large language diffusion model with the Mixture-of-Experts (MoE) architecture, trained from scratch on approximately 20T tokens. LLaDA-MoE achieves compet…
LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion Models
Fengqi Zhu, Rongzhen Wang, Shen Nie +8
While Masked Diffusion Models (MDMs), such as LLaDA, present a promising paradigm for language modeling, there has been relatively little effort in aligning these models with human…
Large Language Diffusion Models
Shen Nie, Fengqi Zhu, Zebin You +7
The capabilities of large language models (LLMs) are widely regarded as relying on autoregressive models (ARMs). We challenge this notion by introducing LLaDA, a diffusion model tr…