4 citations · 4 across the 3 of their papers we have counts for
3 papers
cs.CL2025
LLaDA-MoE: A Sparse MoE Diffusion Language Model
Fengqi Zhu, Zebin You, Yipeng Xing +23
We introduce LLaDA-MoE, a large language diffusion model with the Mixture-of-Experts (MoE) architecture, trained from scratch on approximately 20T tokens. LLaDA-MoE achieves compet…
cs.CL2025
Grove MoE: Towards Efficient and Superior MoE LLMs with Adjugate Experts
Haoyuan Wu, Haoxing Chen, Xiaodong Chen +10
The Mixture of Experts (MoE) architecture is a cornerstone of modern state-of-the-art (SOTA) large language models (LLMs). MoE models facilitate scalability by enabling sparse para…
cs.CV2023★ 4 cited
Few-shot Action Recognition via Intra- and Inter-Video Information Maximization
Huabin Liu, Weiyao Lin, Tieyuan Chen +3
Current few-shot action recognition involves two primary sources of information for classification:(1) intra-video information, determined by frame content within a single video cl…