22 citations · 29 across the 2 of their papers we have counts for
4 papers
Approximate Quantiles for Datacenter Telemetry Monitoring
Gangmuk Lim, Mohamed Hassan, Ze Jin +2
Datacenter systems require efficient troubleshooting and effective resource scheduling so as to minimize downtimes and to efficiently utilize limited resources. In doing so, datace…
StreamBox-HBM: Stream Analytics on High Bandwidth Hybrid Memory
Hongyu Miao, Myeongjae Jeon, Gennady Pekhimenko +2
Stream analytics have an insatiable demand for memory and performance. Emerging hybrid memories combine commodity DDR4 DRAM with 3D-stacked High Bandwidth Memory (HBM) DRAM to meet…
Accelerated Training for CNN Distributed Deep Learning through Automatic Resource-Aware Layer Placement
Jay H. Park, Sunghwan Kim, Jinwon Lee +2
The Convolutional Neural Network (CNN) model, often used for image classification, requires significant training time to obtain high accuracy. To this end, distributed training is…
Analysis of Large-Scale Multi-Tenant GPU Clusters for DNN Training Workloads
Myeongjae Jeon, Shivaram Venkataraman, Amar Phanishayee +3
With widespread advances in machine learning, a number of large enterprises are beginning to incorporate machine learning models across a number of products. These models are typic…