1 paper
Weilin Wan, Jingtao Han, Weizhong Zhang +1
Scaling laws for Large Language Models govern macroscopic resource allocation, yet translating them into precise Mixture-of-Experts (MoE) architectural configurations remains an op…