8 papers · 1 filter
dMoE: dLLMs with Learnable Block Experts
Sicheng Feng, Zigeng Chen, Gongfan Fang +2
Diffusion Large Language Models (dLLMs) have recently emerged as a promising alternative to autoregressive models, offering competitive performance while naturally supporting paral…
Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs
Haiquan Lu, Zigeng Chen, Gongfan Fang +2
LLM agents have recently emerged as a powerful paradigm for solving complex tasks through planning, tool use, memory retrieval, and multi-step interaction. However, these agentic w…
dVoting: Fast Voting for dLLMs
Sicheng Feng, Zigeng Chen, Xinyin Ma +2
Diffusion Large Language Models (dLLMs) represent a new paradigm beyond autoregressive modeling, offering competitive performance while naturally enabling a flexible decoding proce…
dParallel: Learnable Parallel Decoding for dLLMs
Zigeng Chen, Gongfan Fang, Xinyin Ma +2
Diffusion large language models (dLLMs) have recently drawn considerable attention within the research community as a promising alternative to autoregressive generation, offering p…
Efficient Reasoning Models: A Survey
Sicheng Feng, Gongfan Fang, Xinyin Ma +1
Reasoning models have demonstrated remarkable progress in solving complex and logic-intensive tasks by generating extended Chain-of-Thoughts (CoTs) prior to arriving at a final ans…
SparseD: Sparse Attention for Diffusion Language Models
Zeqing Wang, Gongfan Fang, Xinyin Ma +2
While diffusion language models (DLMs) offer a promising alternative to autoregressive models (ARs), existing open-source DLMs suffer from high inference latency. This bottleneck i…