1 paper
Jeremy Herbst, Stefan Wermter, Jae Hee Lee
Mixture-of-Experts (MoE) architectures have become the dominant choice for scaling Large Language Models (LLMs), activating only a subset of parameters per token. While MoE archite…