1 paper
Aaron Archer, Matthew Fahrbach, Kuikui Liu +1
We optimize pipeline parallelism for deep neural network (DNN) inference by partitioning model graphs into k stages and minimizing the running time of the bottleneck stage, inclu…