1 paper
Zongfang Liu, Shengkun Tang, Boyang Sun +2
Sparse Mixture-of-Experts (SMoE) language models achieve strong capability at low per-token compute, yet deployment remains constrained by memory footprint and throughput because t…