14 citations · 31 across the 7 of their papers we have counts for
7 papers
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving
Muhammad Adnan, Rohan Mahapatra, Prashant J. Nair +4
The reasoning and agentic capabilities of large language models have expanded the range of applications they support, from short interactive exchanges to long, compute-heavy reques…
Cross-Model KV Cache Transfer in LLM Families: A Closed-Form Linear Mapping for Prefill Reuse
Taekyung Heo, Rasoul Shafipour, Ritchie Zhao +6
Production deployments often swap between different-sized models in a family for cost-quality cascading, mid-conversation switching, and routing, and each swap forces the receiver…
My CXL Pool Obviates Your PCIe Switch
Yuhong Zhong, Daniel S. Berger, Pantea Zardoshti +5
Pooling PCIe devices across multiple hosts offers a promising solution to mitigate stranded I/O resources, enhance device utilization, address device failures, and reduce total cos…
Coach: Exploiting Temporal Patterns for All-Resource Oversubscription in Cloud Platforms
Benjamin Reidys, Pantea Zardoshti, Íñigo Goiri +16
Cloud platforms remain underutilized despite multiple proposals to improve their utilization (e.g., disaggregation, harvesting, and oversubscription). Our characterization of the r…
Octopus: Enhancing CXL Memory Pods via Sparse Topology
Yuhong Zhong, Fiodar Kazhamiaka, Pantea Zardoshti +4
The Compute Express Link (CXL) interconnect enables compute "pods" that pool memory across servers to reduce cost and improve efficiency. These pods also facilitate pairwise commun…
Workload Intelligence: Punching Holes Through the Cloud Abstraction
Lexiang Huang, Anjaly Parayil, Jue Zhang +13
Today, cloud workloads are essentially opaque to the cloud platform. Typically, the only information the platform receives is the virtual machine (VM) type and possibly a decoratio…