2 papers
cs.DC2023
Load is not what you should balance: Introducing Prequal
Bartek Wydrowski, Robert Kleinberg, Stephen M. Rumble +1
We present Prequal (Probing to Reduce Queuing and Latency), a load balancer for distributed multi-tenant systems. Prequal aims to minimize real-time request latency in the presence…
cs.LG2023
Practical Performance Guarantees for Pipelined DNN Inference
Aaron Archer, Matthew Fahrbach, Kuikui Liu +1
We optimize pipeline parallelism for deep neural network (DNN) inference by partitioning model graphs into stages and minimizing the running time of the bottleneck stage, inclu…