1 paper · 1 filter
Minyu Cui, Anna Wingkvist, Morgan Ericsson
Mixture-of-Experts (MoE) architectures increase model capacity without proportionally increasing computation cost and have become a key building block for scaling large language mo…