activity
20212026
most citedInference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

6 citations · 9 across the 9 of their papers we have counts for

collaborators
Showing cs.DCShow all

6 papers · 1 filter

cs.DC2026

StrataCL: Fabric-Native Communication Library for Production Supernodes

Tiancheng Hu, Jin Qin, Yuzheng Wang +14

Modern distributed AI workloads run across hundreds of accelerators, making communication a major bottleneck. Existing communication libraries remain largely buffer-centric because…

cs.DC2026

Tessera: Unlocking Heterogeneous GPUs through Kernel-Granularity Disaggregation

Tiancheng Hu, Jin Qin, Zheng Wang +10

Disaggregation maps parts of an AI workload to different types of GPUs, offering a path to utilize modern heterogeneous GPU clusters. However, existing solutions operate at a coars…

cs.DC2026

Hummingbird: SLO-Oriented GPU Preemption at Microsecond-scale

Tiancheng Hu, Chenxi Wang, Ting Cao +9

Existing GPU-sharing techniques, including spatial and temporal sharing, aim to improve utilization but face challenges in simultaneously ensuring SLO adherence and maximizing effi…

cs.DC2024

A Tale of Two Paths: Toward a Hybrid Data Plane for Efficient Far-Memory Applications

Lei Chen, Shi Liu, Chenxi Wang +9

With rapid advances in network hardware, far memory has gained a great deal of traction due to its ability to break the memory capacity wall. Existing far memory systems fall into…

cs.DC20243 cited

MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Cunchen Hu, Heyang Huang, Junhao Hu +8

Large language model (LLM) serving has transformed from stateless to stateful systems, utilizing techniques like context caching and disaggregated inference. These optimizations ex…

cs.DC20246 cited

Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Cunchen Hu, Heyang Huang, Liangliang Xu +9

Transformer-based large language model (LLM) inference serving is now the backbone of many cloud services. LLM inference consists of a prefill phase and a decode phase. However, ex…