1 paper · 1 filter
Ruisi Cai, Yeonju Ro, Geon-Woo Kim +4
The proliferation of large language models (LLMs) has led to the adoption of Mixture-of-Experts (MoE) architectures that dynamically leverage specialized subnetworks for improved e…