19 citations · 59 across the 14 of their papers we have counts for
7 papers · 1 filter
Dirigo: Self-scaling Stateful Actors For Serverless Real-time Data Processing
Le Xu, Divyanshu Saxena, Neeraja J. Yadwadkar +2
We propose Dirigo, a distributed stream processing service built atop virtual actors. Dirigo achieves both a high level of resource efficiency and performance isolation driven by u…
Shockwave: Fair and Efficient Cluster Scheduling for Dynamic Adaptation in Machine Learning
Pengfei Zheng, Rui Pan, Tarannum Khan +2
Dynamic adaptation has become an essential technique in accelerating distributed machine learning (ML) training. Recent studies have shown that dynamically adjusting model structur…
Elastic Model Aggregation with Parameter Service
Juncheng Gu, Mosharaf Chowdhury, Kang G. Shin +1
Model aggregation, the process that updates model parameters, is an important step for model convergence in distributed deep learning (DDL). However, the parameter server (PS), a p…
Archipelago: A Scalable Low-Latency Serverless Platform
Arjun Singhvi, Kevin Houck, Arjun Balasubramanian +3
The increased use of micro-services to build web applications has spurred the rapid growth of Function-as-a-Service (FaaS) or serverless computing platforms. While FaaS simplifies…
SNF: Serverless Network Functions
Arjun Singhvi, Junaid Khalid, Aditya Akella +1
It is increasingly common to outsource network functions (NFs) to the cloud. However, no cloud providers offer NFs-as-a-Service (NFaaS) that allows users to run custom NFs. Our wor…
Themis: Fair and Efficient GPU Cluster Scheduling
Kshiteej Mahajan, Arjun Balasubramanian, Arjun Singhvi +4
Modern distributed machine learning (ML) training workloads benefit significantly from leveraging GPUs. However, significant contention ensues when multiple such workloads are run…