2 papers
cs.CL2026
Rethinking Token Prediction: Tree-Structured Diffusion Language Model
Zihao Wu, Haoming Yang, Juncheng Dong +1
Discrete diffusion language models have emerged as a competitive alternative to auto-regressive language models, but training them efficiently under limited parameter and memory bu…
cs.LG2025
Input Domain Aware MoE: Decoupling Routing Decisions from Task Optimization in Mixture of Experts
Yongxiang Hua, Haoyu Cao, Zhou Tao +4
Sparse Mixture of Experts (sMoE) has become a pivotal approach for scaling large vision-language models, offering substantial capacity while maintaining computational efficiency th…