27 citations · 38 across the 5 of their papers we have counts for
3 papers · 1 filter
Train Where the Data is: A Case for Bandwidth Efficient Coded Training
Zhifeng Lin, Krishna Giri Narra, Mingchao Yu +2
Training a machine learning model is both compute and data-intensive. Most of the model training is performed on high performance compute nodes and the training data is stored near…
Collage Inference: Achieving low tail latency during distributed image classification using coded redundancy models
Krishna Narra, Zhifeng Lin, Ganesh Ananthanarayanan +2
Reducing the latency variance in machine learning inference is a key requirement in many applications. Variance is harder to control in a cloud deployment in the presence of stragg…
Slack Squeeze Coded Computing for Adaptive Straggler Mitigation
Krishna Giri Narra, Zhifeng Lin, Mehrdad Kiamari +2
While performing distributed computations in today's cloud-based platforms, execution speed variations among compute nodes can significantly reduce the performance and create bottl…