3 citations · 3 across the 1 of their papers we have counts for
Showing cs.DCShow all
3 papers · 1 filter
cs.DC2024
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
Jiabin Chen, Fei Xu, Yikun Gu +3
Deep Neural Network (DNN) inference on serverless functions is gaining prominence due to its potential for substantial budget savings. Existing works on serverless DNN inference so…
cs.DC2023
Opara: Exploiting Operator Parallelism for Expediting DNN Inference on GPUs
Aodong Chen, Fei Xu, Li Han +4
GPUs have become the \emph{defacto} hardware devices for accelerating Deep Neural Network (DNN) inference workloads. However, the conventional \emph{sequential execution mode of DN…
cs.DC2022★ 3 cited
iGniter: Interference-Aware GPU Resource Provisioning for Predictable DNN Inference in the Cloud
Fei Xu, Jianian Xu, Jiabin Chen +4
GPUs are essential to accelerating the latency-sensitive deep neural network (DNN) inference workloads in cloud datacenters. To fully utilize GPU resources, spatial sharing of GPUs…