3 papers
cs.DC2026
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving
Muhammad Adnan, Rohan Mahapatra, Prashant J. Nair +4
The reasoning and agentic capabilities of large language models have expanded the range of applications they support, from short interactive exchanges to long, compute-heavy reques…
cs.AR2026
Octopus: Enhancing CXL Memory Pods via Sparse Topology
Yuhong Zhong, Fiodar Kazhamiaka, Pantea Zardoshti +4
The Compute Express Link (CXL) interconnect enables compute "pods" that pool memory across servers to reduce cost and improve efficiency. These pods also facilitate pairwise commun…
cs.OS2025
Coach: Exploiting Temporal Patterns for All-Resource Oversubscription in Cloud Platforms
Benjamin Reidys, Pantea Zardoshti, Ãñigo Goiri +16
Cloud platforms remain underutilized despite multiple proposals to improve their utilization (e.g., disaggregation, harvesting, and oversubscription). Our characterization of the r…