collaborators

6 papers

cs.AI2026

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models

Fengqi Zhu, Shaoxuan Xu, Jingyang Ou +11

Diffusion language models (dLLMs) offer an alternative to autoregressive (AR) language modeling, yet the scaling behavior of Mixture-of-Experts (MoE) dLLMs remains poorly understoo…

cs.CL2026

Improved Large Language Diffusion Models

Shen Nie, Qiyang Min, Shaoxuan Xu +7

Modern large language models are predominantly trained with autoregressive factorization and causal attention. We present \emph{iLLaDA}, an 8B masked diffusion language model train…

cs.CV2026

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs

Yaoting Wang, Ziyi Zhang, Wenming Tu +10

Recent advances in Omni-Multimodal Large Language Models (Omni-MLLMs) have enabled strong integration of vision, audio, and language. However, their audio-visual intelligence (AVI)…

cs.AI2026

DLLM-Searcher: Adapting Diffusion Large Language Model for Search Agents

Jiahao Zhao, Shaoxuan Xu, Zhongxiang Sun +6

Recently, Diffusion Large Language Models (dLLMs) have demonstrated unique efficiency advantages, enabled by their inherently parallel decoding mechanism and flexible generation pa…

cs.CL2025

Principled RL for Diffusion LLMs Emerges from a Sequence-Level Perspective

Jingyang Ou, Jiaqi Han, Minkai Xu +5

Reinforcement Learning (RL) has proven highly effective for autoregressive language models, but adapting these methods to diffusion large language models (dLLMs) presents fundament…

cs.LG2025

BalanceBenchmark: A Survey for Multimodal Imbalance Learning

Shaoxuan Xu, Menglu Cui, Chengxiang Huang +2

Multimodal learning has gained attention for its capacity to integrate information from different modalities. However, it is often hindered by the multimodal imbalance problem, whe…