1 paper
Munkyu Lee, Sihoon Seong, Minki Kang +5
In cloud environments, GPU-based deep neural network (DNN) inference servers are required to meet the Service Level Objective (SLO) latency for each workload under a specified requ…