1 paper
Yuchen Yang, Yaru Zhao, Pu Yang +2
While Mixture-of-Experts (MoE) architectures substantially bolster the expressive power of large-language models, their prohibitive memory footprint severely impedes the practical…