131 citations · 138 across the 4 of their papers we have counts for
8 papers
A Simulation Platform for Multi-tenant Machine Learning Services on Thousands of GPUs
Ruofan Liang, Bingsheng He, Shengen Yan +1
Multi-tenant machine learning services have become emerging data-intensive workloads in data centers with heavy usage of GPU resources. Due to the large scale, many tuning paramete…
Characterization and Prediction of Deep Learning Workloads in Large-Scale GPU Datacenters
Qinghao Hu, Peng Sun, Shengen Yan +2
Modern GPU datacenters are critical for delivering Deep Learning (DL) models and services in both the research community and industry. When operating a datacenter, optimization of…
ModelCI-e: Enabling Continual Learning in Deep Learning Serving Systems
Yizheng Huang, Huaizheng Zhang, Yonggang Wen +2
MLOps is about taking experimental ML models to production, i.e., serving the models to actual users. Unfortunately, existing ML serving systems do not adequately handle the dynami…
Optimizing Network Performance for Distributed DNN Training on GPU Clusters: ImageNet/AlexNet Training in 1.5 Minutes
Peng Sun, Wansen Feng, Ruobing Han +2
It is important to scale out deep neural network (DNN) training for reducing model training time. The high communication overhead is one of the major performance bottlenecks for di…
GraphMP: I/O-Efficient Big Graph Analytics on a Single Commodity Machine
Peng Sun, Yonggang Wen, Ta Nguyen Binh Duong +1
Recent studies showed that single-machine graph processing systems can be as highly competitive as cluster-based approaches on large-scale problems. While several out-of-core graph…
Speeding-up Age Estimation in Intelligent Demographics System via Network Optimization
Zhenzhen Hui, Peng Sun, Yonggang Wen
Age estimation is a difficult task which requires the automatic detection and interpretation of facial features. Recently, Convolutional Neural Networks (CNNs) have made remarkable…