1 paper · 1 filter
Kexin Chu, Dawei Xiang, Zixu Shen +3
Mixture-of-Experts (MoE) has become a practical architecture for scaling LLM capacity while keeping per-token compute modest, but deploying MoE models on a single, memory-limited G…