3 papers
cs.LG2026
ToMoE: Converting Dense Large Language Models to Mixture-of-Experts through Dynamic Structural Pruning
Shangqian Gao, Ting Hua, Reza Shirkavand +10
Large Language Models (LLMs) have demonstrated remarkable abilities in tackling a wide range of complex tasks. However, their huge computational and memory costs raise significant…
cs.CL2025
MossNet: Mixture of State-Space Experts is a Multi-Head Attention
Shikhar Tuli, James Seale Smith, Haris Jeelani +5
Large language models (LLMs) have significantly advanced generative applications in natural language processing (NLP). Recent trends in model architectures revolve around efficient…
cs.LG2025
Retraining-Free Merging of Sparse MoE via Hierarchical Clustering
I-Chun Chen, Hsu-Shen Liu, Wei-Fang Sun +3
Sparse Mixture-of-Experts (SMoE) models represent a significant advancement in large language model (LLM) development through their efficient parameter utilization. These models ac…