1 paper
Taehyun Kim, Kwanseok Choi, Youngmock Cho +3
Mixture-of-Experts (MoE) large language models (LLM) have memory requirements that often exceed the GPU memory capacity, requiring costly parameter movement from secondary memories…