1 paper · 1 filter
Alfarizy Alfarizy, Hung Truong Thanh Nguyen, René Richard +2
Mixture-of-Experts (MoE) language models are often described as ideal for resource-constrained inference. Each token activates only a small subset of experts, so the per-token comp…