3 papers
cs.LG2026
ReMoE: Boosting Expert Reuse through Router Fine-Tuning in Memory-Constrained MoE LLM Inference
Xiongwei Zhu, Xiaojian Liao, Tianyang Jiang +3
Fine-grained Mixture-of-Experts (MoE) models sparsely activate only a subset of experts per token, reducing activated computation while maintaining high model capacity. However, in…
cs.AR2025
CaMDN: Enhancing Cache Efficiency for Multi-tenant DNNs on Integrated NPUs
Tianhao Cai, Liang Wang, Limin Xiao +4
With the rapid development of DNN applications, multi-tenant execution, where multiple DNNs are co-located on a single SoC, is becoming a prevailing trend. Although many methods ar…
cs.DC2025
PipeBoost: Resilient Pipelined Architecture for Fast Serverless LLM Scaling
Chongpeng Liu, Xiaojian Liao, Hancheng Liu +2
This paper presents PipeBoost, a low-latency LLM serving system for multi-GPU (serverless) clusters, which can rapidly launch inference services in response to bursty requests with…