27 citations · 38 across the 4 of their papers we have counts for
6 papers
Privacy-Preserving Inference in Machine Learning Services Using Trusted Execution Environments
Krishna Giri Narra, Zhifeng Lin, Yongqin Wang +2
This work presents Origami, which provides privacy-preserving inference for large deep neural network (DNN) models through a combination of enclave execution, cryptographic blindin…
Train Where the Data is: A Case for Bandwidth Efficient Coded Training
Zhifeng Lin, Krishna Giri Narra, Mingchao Yu +2
Training a machine learning model is both compute and data-intensive. Most of the model training is performed on high performance compute nodes and the training data is stored near…
Collage Inference: Achieving low tail latency during distributed image classification using coded redundancy models
Krishna Narra, Zhifeng Lin, Ganesh Ananthanarayanan +2
Reducing the latency variance in machine learning inference is a key requirement in many applications. Variance is harder to control in a cloud deployment in the presence of stragg…
Collage Inference: Using Coded Redundancy for Low Variance Distributed Image Classification
Krishna Giri Narra, Zhifeng Lin, Ganesh Ananthanarayanan +2
MLaaS (ML-as-a-Service) offerings by cloud computing platforms are becoming increasingly popular. Hosting pre-trained machine learning models in the cloud enables elastic scalabili…
Slack Squeeze Coded Computing for Adaptive Straggler Mitigation
Krishna Giri Narra, Zhifeng Lin, Mehrdad Kiamari +2
While performing distributed computations in today's cloud-based platforms, execution speed variations among compute nodes can significantly reduce the performance and create bottl…
GradiVeQ: Vector Quantization for Bandwidth-Efficient Gradient Aggregation in Distributed CNN Training
Mingchao Yu, Zhifeng Lin, Krishna Narra +6
Data parallelism can boost the training speed of convolutional neural networks (CNN), but could suffer from significant communication costs caused by gradient aggregation. To allev…