collaborators

5 papers

cs.DC2026

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving

Muhammad Adnan, Rohan Mahapatra, Prashant J. Nair +4

The reasoning and agentic capabilities of large language models have expanded the range of applications they support, from short interactive exchanges to long, compute-heavy reques…

cs.LG2026

Cross-Model KV Cache Transfer in LLM Families: A Closed-Form Linear Mapping for Prefill Reuse

Taekyung Heo, Rasoul Shafipour, Ritchie Zhao +6

Production deployments often swap between different-sized models in a family for cost-quality cascading, mid-conversation switching, and routing, and each swap forces the receiver…

cs.AR2026

Octopus: Enhancing CXL Memory Pods via Sparse Topology

Yuhong Zhong, Fiodar Kazhamiaka, Pantea Zardoshti +4

The Compute Express Link (CXL) interconnect enables compute "pods" that pool memory across servers to reduce cost and improve efficiency. These pods also facilitate pairwise commun…

cs.OS2025

My CXL Pool Obviates Your PCIe Switch

Yuhong Zhong, Daniel S. Berger, Pantea Zardoshti +5

Pooling PCIe devices across multiple hosts offers a promising solution to mitigate stranded I/O resources, enhance device utilization, address device failures, and reduce total cos…

cs.OS2025

Coach: Exploiting Temporal Patterns for All-Resource Oversubscription in Cloud Platforms

Benjamin Reidys, Pantea Zardoshti, Íñigo Goiri +16

Cloud platforms remain underutilized despite multiple proposals to improve their utilization (e.g., disaggregation, harvesting, and oversubscription). Our characterization of the r…