most citedPrivacy-Preserving Inference in Machine Learning Services Using Trusted Execution Environments

27 citations · 38 across the 4 of their papers we have counts for

collaborators

6 papers

cs.LG201927 cited

Privacy-Preserving Inference in Machine Learning Services Using Trusted Execution Environments

Krishna Giri Narra, Zhifeng Lin, Yongqin Wang +2

This work presents Origami, which provides privacy-preserving inference for large deep neural network (DNN) models through a combination of enclave execution, cryptographic blindin…

cs.DC20191 cited

Train Where the Data is: A Case for Bandwidth Efficient Coded Training

Zhifeng Lin, Krishna Giri Narra, Mingchao Yu +2

Training a machine learning model is both compute and data-intensive. Most of the model training is performed on high performance compute nodes and the training data is stored near…

cs.DC2019

Collage Inference: Achieving low tail latency during distributed image classification using coded redundancy models

Krishna Narra, Zhifeng Lin, Ganesh Ananthanarayanan +2

Reducing the latency variance in machine learning inference is a key requirement in many applications. Variance is harder to control in a cloud deployment in the presence of stragg…

cs.CV2019

Collage Inference: Using Coded Redundancy for Low Variance Distributed Image Classification

Krishna Giri Narra, Zhifeng Lin, Ganesh Ananthanarayanan +2

MLaaS (ML-as-a-Service) offerings by cloud computing platforms are becoming increasingly popular. Hosting pre-trained machine learning models in the cloud enables elastic scalabili…

cs.DC2019

Slack Squeeze Coded Computing for Adaptive Straggler Mitigation

Krishna Giri Narra, Zhifeng Lin, Mehrdad Kiamari +2

While performing distributed computations in today's cloud-based platforms, execution speed variations among compute nodes can significantly reduce the performance and create bottl…

cs.LG201810 cited

GradiVeQ: Vector Quantization for Bandwidth-Efficient Gradient Aggregation in Distributed CNN Training

Mingchao Yu, Zhifeng Lin, Krishna Narra +6

Data parallelism can boost the training speed of convolutional neural networks (CNN), but could suffer from significant communication costs caused by gradient aggregation. To allev…