activity
20242026
collaborators

7 papers

cs.AI2026

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models

Fengqi Zhu, Shaoxuan Xu, Jingyang Ou +11

Diffusion language models (dLLMs) offer an alternative to autoregressive (AR) language modeling, yet the scaling behavior of Mixture-of-Experts (MoE) dLLMs remains poorly understoo…

cs.CL2026

Improved Large Language Diffusion Models

Shen Nie, Qiyang Min, Shaoxuan Xu +7

Modern large language models are predominantly trained with autoregressive factorization and causal attention. We present \emph{iLLaDA}, an 8B masked diffusion language model train…

cs.CL2025

Large Language Diffusion Models

Shen Nie, Fengqi Zhu, Zebin You +7

The capabilities of large language models (LLMs) are widely regarded as relying on autoregressive models (ARMs). We challenge this notion by introducing LLaDA, a diffusion model tr…

cs.LG2025

LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion Models

Fengqi Zhu, Rongzhen Wang, Shen Nie +8

While Masked Diffusion Models (MDMs), such as LLaDA, present a promising paradigm for language modeling, there has been relatively little effort in aligning these models with human…

cs.MA2025

GenSim: A General Social Simulation Platform with Large Language Model based Agents

Jiakai Tang, Heyang Gao, Xuchen Pan +11

With the rapid advancement of large language models (LLMs), recent years have witnessed many promising studies on leveraging LLM-based agents to simulate human social behavior. Whi…

cs.CL2025

DeepCritic: Deliberate Critique with Large Language Models

Wenkai Yang, Jingwen Chen, Yankai Lin +1

As Large Language Models (LLMs) are rapidly evolving, providing accurate feedback and scalable oversight on their outputs becomes an urgent and critical problem. Leveraging LLMs as…