4 papers
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving
Xingqi Cui, Chieh-Jan Mike Liang, Ziang Tang +2
Achieving cost efficiency while meeting strict user-facing SLOs (e.g., time-to-first-token) remains a fundamental challenge for cloud GPU clusters serving large language models (LL…
From Models to Operators: Rethinking Autoscaling Granularity for Large Generative Models
Xingqi Cui, Chieh-Jan Mike Liang, Jiarong Xing +1
Serving large generative models such as LLMs and multi- modal transformers requires balancing user-facing SLOs (e.g., time-to-first-token, time-between-tokens) with provider goals…
MetaMuse: Algorithm Generation via Creative Ideation
Ruiying Ma, Chieh-Jan Mike Liang, Yanjie Gao +1
Designing system algorithms remains challenging, where the discontinuous nature of the solution space often forces system engineers to rely on generic heuristics at the expense of…
DEX: Scalable Range Indexing on Disaggregated Memory [Extended Version]
Baotong Lu, Kaisong Huang, Chieh-Jan Mike Liang +2
Memory disaggregation can potentially allow memory-optimized range indexes such as B+-trees to scale beyond one machine while attaining high hardware utilization and low cost. Desi…