1 paper
Seeyeon Kim, Juhyeong Jin, Joo-Young Kim
Modern mixture-of-experts (MoE) language models increasingly strain the capacity and cost efficiency of high-bandwidth memory (HBM), as rapidly growing expert weights must be provi…