2 papers
cs.DC2022
PARIS and ELSA: An Elastic Scheduling Algorithm for Reconfigurable Multi-GPU Inference Servers
Yunseong Kim, Yujeong Choi, Minsoo Rhu
In cloud machine learning (ML) inference systems, providing low latency to end-users is of utmost importance. However, maximizing server utilization and system throughput is also c…
cs.DC2020
LazyBatching: An SLA-aware Batching System for Cloud Machine Learning Inference
Yujeong Choi, Yunseong Kim, Minsoo Rhu
In cloud ML inference systems, batching is an essential technique to increase throughput which helps optimize total-cost-of-ownership. Prior graph batching combines the individual…