1 paper · 1 filter
Seungwoo Jung, Dohyeok Kwon, Seungmin Cha +4
Mixture-of-experts (MoE) architectures scale large language models efficiently, but they demand massive GPU memory. To cope with such demand, models are commonly compressed to redu…