4 citations · 4 across the 1 of their papers we have counts for
2 papers
cs.DC2022
A Survey of Multi-Tenant Deep Learning Inference on GPU
Fuxun Yu, Di Wang, Longfei Shangguan +3
Deep Learning (DL) models have achieved superior performance. Meanwhile, computing hardware like NVIDIA GPUs also demonstrated strong computing scaling trends with 2x throughput an…
cs.AR2020★ 4 cited
Towards Latency-aware DNN Optimization with GPU Runtime Analysis and Tail Effect Elimination
Fuxun Yu, Zirui Xu, Tong Shen +12
Despite the superb performance of State-Of-The-Art (SOTA) DNNs, the increasing computational cost makes them very challenging to meet real-time latency and accuracy requirements. A…