1 paper
Qianchao Zhu, Xucheng Ye, Yuliang Liu +2
Mixture-of-Experts models have become a dominant architecture for scaling Large Language Models by activating only a sparse subset of experts per token. However, latency-critical M…