1 paper
Cedric Chan, Jona te Lintelo, Stjepan Picek
Mixture of Experts (MoE) architectures have gained popularity for reducing computational costs in deep neural networks by activating only a subset of parameters during inference. W…