activity
20242026
collaborators

8 papers

cs.DC2026

Huawei Cloud Model-as-a-Service on the CloudMatrix384 SuperPod

Ao Xiao, Bangzheng He, Baoquan Zhang +125

Scaled-out MoE LLMs and scaled-up SuperPods create new systems challenges for production Model-as-a-Service (MaaS), requiring disaggregation, low-latency communication, and decentr…

cs.DC2025

Efficient Serving of LLM Applications with Probabilistic Demand Modeling

Yifei Liu, Zuo Gan, Zhenghao Gan +8

Applications based on Large Language Models (LLMs) contains a series of tasks to address real-world problems with boosted capability, which have dynamic demand volumes on diverse b…

cs.DC2025

DeepServe: Serverless Large Language Model Serving at Scale

Junhao Hu, Jiang Xu, Zhixia Liu +18

In this paper, we propose DEEPSERVE, a scalable and serverless AI platform designed to efficiently serve large language models (LLMs) at scale in cloud environments. DEEPSERVE addr…

cs.LG2025

RaaS: Reasoning-Aware Attention Sparsity for Efficient LLM Reasoning

Junhao Hu, Wenrui Huang, Weidong Wang +6

Large Language Models (LLMs) have demonstrated strong capabilities across various domains, with recent advancements in challenging reasoning tasks such as mathematics and programmi…

cs.LG2025

EPIC: Efficient Position-Independent Caching for Serving Large Language Models

Junhao Hu, Wenrui Huang, Weidong Wang +7

Large Language Models (LLMs) show great capabilities in a wide range of applications, but serving them efficiently becomes increasingly challenging as requests (prompts) become mor…

cs.DB2025

K2: On Optimizing Distributed Transactions in a Multi-region Data Store with TrueTime Clocks (Extended Version)

Haoze Song, Yongqi Wang, Xusheng Chen +5

TrueTime clocks (TTCs) that offer accurate and reliable time within limited uncertainty bounds have been increasingly implemented in many clouds. Multi-region data stores that seek…