1 citations · 1 across the 1 of their papers we have counts for
Showing cs.DCShow all
3 papers · 1 filter
cs.DC2025
C-Koordinator: Interference-aware Management for Large-scale and Co-located Microservice Clusters
Shengye Song, Minxian Xu, Zuowei Zhang +5
Microservices transform traditional monolithic applications into lightweight, loosely coupled application components and have been widely adopted in many enterprises. Cloud platfor…
cs.DC2025★ 1 cited
Serving Large Language Models on Huawei CloudMatrix384
Pengfei Zuo, Huimin Lin, Junbo Deng +43
The rapid evolution of large language models (LLMs), driven by growing parameter scales, adoption of mixture-of-experts (MoE) architectures, and expanding context lengths, imposes…
cs.DC2025
HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public Clouds
Chiheng Lou, Sheng Qi, Chao Jin +5
With the proliferation of large language model (LLM) variants, developers are turning to serverless computing for cost-efficient LLM deployment. However, public cloud providers oft…