10 citations · 16 across the 4 of their papers we have counts for
4 papers · 1 filter
Cloud Collectives: Towards Cloud-aware Collectives forML Workloads with Rank Reordering
Liang Luo, Jacob Nelson, Arvind Krishnamurthy +1
ML workloads are becoming increasingly popular in the cloud. Good cloud training performance is contingent on efficient parameter exchange among VMs. We find that Collectives, the…
Scaling Distributed Machine Learning with In-Network Aggregation
Amedeo Sapio, Marco Canini, Chen-Yu Ho +7
Training machine learning models in parallel is an increasingly important workload. We accelerate distributed parallel training by designing a communication primitive that uses a p…
ADARES: Adaptive Resource Management for Virtual Machines
Ignacio Cano, Lequn Chen, Pedro Fonseca +5
Virtual execution environments allow for consolidation of multiple applications onto the same physical server, thereby enabling more efficient use of server resources. However, use…
Parameter Hub: a Rack-Scale Parameter Server for Distributed Deep Neural Network Training
Liang Luo, Jacob Nelson, Luis Ceze +2
Distributed deep neural network (DDNN) training constitutes an increasingly important workload that frequently runs in the cloud. Larger DNN models and faster compute engines are s…