1 paper · 1 filter
Zhenyu Liu, Yunzhen Liu, Zehao Fan +5
Mixture-of-Experts (MoE) models scale capacity via sparse activation but stress memory and bandwidth. Offloading alleviates GPU memory by fetching experts on demand, yet token-leve…