1 paper
Daniel Fidel Harvey, George Weale, Berk Yilmaz
Mixture of Experts (MoE) architectures increase large language model scalability, yet their performance depends on the router module that moves tokens to specialized experts. Bad r…