1 paper
Zhuoren Ye, Tianyu Wo, Dinghao Xue +4
Emerging LLM services increasingly host many sparse MoE models, yet most models receive sparse requests and remain cold. This creates a GPU memory problem: model weights are stable…