10 citations · 13 across the 3 of their papers we have counts for
3 papers
cs.LG2026
Strait: Perceiving Priority and Interference in ML Inference Serving
Haidong Zhao, Nikolaos Georgantas
Machine learning (ML) inference serving systems host deep neural network (DNN) models and schedule incoming inference requests across deployed GPUs. However, limited support for ta…
cs.LG2025★ 3 cited
ML Inference Scheduling with Predictable Latency
Haidong Zhao, Nikolaos Georgantas
Machine learning (ML) inference serving systems can schedule requests to improve GPU utilization and to meet service level objectives (SLOs) or deadlines. However, improving GPU ut…
cs.DC2022★ 10 cited
Supporting Multi-Cloud in Serverless Computing
Haidong Zhao, Zakaria Benomar, Tobias Pfandzelter +1
Serverless computing is a widely adopted cloud execution model composed of Function-as-a-Service (FaaS) and Backend-as-a-Service (BaaS) offerings. The increased level of abstractio…