activity
20242026
collaborators

10 papers

cs.PF2026

SPLIT: SymPathy for Large jobs Improves Tail latency

Zhouzi Li, Mor Harchol-Balter, Alan Scheller-Wolf

We study the asymptotic response time tail in the M/G/n multi-server queue with heavy-tailed (regularly varying) job sizes, a setting representative of modern computing workloads.…

cs.DC2026

Mean field optimal Core Allocation across Malleable jobs

Zhouzi Li, Mor Harchol-Balter, Benjamin Berg

Modern data centers and cloud computing clusters are increasingly running workloads composed of malleable jobs. A malleable job can be parallelized across any number of cores, yet…

cs.DC2026

BOA Constrictor: Squeezing Performance out of GPUs in the Cloud via Budget-Optimal Allocation

Zhouzi Li, Cindy Zhu, Arpan Mukhopadhyay +2

The past decade has seen a dramatic increase in demand for GPUs to train Machine Learning (ML) models. Because it is prohibitively expensive for most organizations to build and mai…

cs.PF2026

LookAhead: The Optimal Non-decreasing Index Policy for a Time-Varying Holding Cost problem

Keerthana Gurushankar, Zhouzi Li, Mor Harchol-Balter +1

In practice, the cost of delaying a job can grow as the job waits. Such behavior is modeled by the Time-Varying Holding Cost (TVHC) problem, where each job's instantaneous holding…

cs.PF2025

Asymptotically Optimal Scheduling of Multiple Parallelizable Job Classes

Benjamin Berg, Benjamin Moseley, Weina Wang +1

Modern computing workloads are often composed of parallelizable jobs. A parallelizable job can be completed more quickly when run on additional servers. However, each job can only…

cs.PF2025

An Upper Bound on the M/M/k Queue With Deterministic Setup Times

Jalani Williams, Weina Wang, Mor Harchol-Balter

In many systems, servers do not turn on instantly; instead, a setup time must pass before a server can begin work. These "setup times" can wreak havoc on a system's queueing; this…