2 papers
cs.LG2025
Secret mixtures of experts inside your LLM
Enric Boix-Adsera
Despite being one of the earliest neural network layers, the Multilayer Perceptron (MLP) is arguably one of the least understood parts of the transformer architecture due to its de…
cs.LG2025
The power of fine-grained experts: Granularity boosts expressivity in Mixture of Experts
Enric Boix-Adsera, Philippe Rigollet
Mixture-of-Experts (MoE) layers are increasingly central to frontier model architectures. By selectively activating parameters, they reduce computational cost while scaling total p…