25 citations · 157 across the 23 of their papers we have counts for
6 papers · 1 filter
TEMPI: An Interposed MPI Library with a Canonical Representation of CUDA-aware Datatypes
Carl Pearson, Kun Wu, I-Hsin Chung +2
MPI derived datatypes are an abstraction that simplifies handling of non-contiguous data in MPI applications. These datatypes are recursively constructed at runtime from primitive…
At-Scale Sparse Deep Neural Network Inference with Efficient GPU Implementation
Mert Hidayetoglu, Carl Pearson, Vikram Sharma Mailthody +4
This paper presents GPU performance optimization and scaling results for inference models of the Sparse Deep Neural Network Challenge 2020. Demands for network quality have increas…
EMOGI: Efficient Memory-access for Out-of-memory Graph-traversal In GPUs
Seung Won Min, Vikram Sharma Mailthody, Zaid Qureshi +3
Modern analytics and recommendation systems are increasingly based on graph data that capture the relations between entities being analyzed. Practical graphs come in huge sizes, of…
MLModelScope: A Distributed Platform for Model Evaluation and Benchmarking at Scale
Abdul Dakkak, Cheng Li, Jinjun Xiong +1
Machine Learning (ML) and Deep Learning (DL) innovations are being introduced at such a rapid pace that researchers are hard-pressed to analyze and study them. The complicated proc…
Helios: Heterogeneity-Aware Federated Learning with Dynamically Balanced Collaboration
Zirui Xu, Fuxun Yu, Jinjun Xiong +1
In this paper, we propose Helios, a heterogeneity-aware FL framework to tackle the straggler issue. Helios identifies individual devices' heterogeneous training capability, and the…
TrIMS: Transparent and Isolated Model Sharing for Low Latency Deep LearningInference in Function as a Service Environments
Abdul Dakkak, Cheng Li, Simon Garcia de Gonzalo +2
Deep neural networks (DNNs) have become core computation components within low latency Function as a Service (FaaS) prediction pipelines: including image recognition, object detect…