1 paper
Margaret Li, Sneha Kudugunta, Danielle Rothermel +1
Mixture-of-Experts (MoE) architectures have become standard in large language models, yet many of their core design choices - expert count, granularity, shared experts, load balanc…