Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
MoE-SpeQ: Speculative Quantized Decoding with Proactive Expert Prefetching and Offloading for Mixture-of-Experts
Wenfeng Wang, Jiacheng Liu, Xiaofeng Hou +5
The immense memory requirements of state-of-the-art Mixture-of-Experts (MoE) models present a significant challenge for inference, often exceeding the capacity of a single accelera…
cs.LG2025★ 9 cited
A Survey on Inference Optimization Techniques for Mixture of Experts Models
Jiacheng Liu, Peng Tang, Wenfeng Wang +5
The emergence of large-scale Mixture of Experts (MoE) models represents a significant advancement in artificial intelligence, offering enhanced model capacity and computational eff…