activity
20162025
most citedAn Efficient Statistical-based Gradient Compression Technique for Distributed Training Systems

31 citations · 73 across the 9 of their papers we have counts for

collaborators
Showing cs.DCShow all

6 papers · 1 filter

cs.DC20213 cited

With Great Freedom Comes Great Opportunity: Rethinking Resource Allocation for Serverless Functions

Muhammad Bilal, Marco Canini, Rodrigo Fonseca +1

Current serverless offerings give users a limited degree of flexibility for configuring the resources allocated to their function invocations by either coupling memory and CPU reso…

cs.DC2019

On the Discrepancy between the Theoretical Analysis and Practical Implementations of Compressed Communication for Distributed Deep Learning

Aritra Dutta, El Houcine Bergou, Ahmed M. Abdelmoniem +4

Compressed communication, in the form of sparsification or quantization of stochastic gradients, is employed to reduce communication costs in distributed data-parallel training of…

cs.DC2019

Assise: Performance and Availability via NVM Colocation in a Distributed File System

Thomas E. Anderson, Marco Canini, Jongyul Kim +6

The adoption of very low latency persistent memory modules (PMMs) upends the long-established model of disaggregated file system access. Instead, by colocating computation and PMM…

cs.DC2019

Scaling Distributed Machine Learning with In-Network Aggregation

Amedeo Sapio, Marco Canini, Chen-Yu Ho +7

Training machine learning models in parallel is an increasingly important workload. We accelerate distributed parallel training by designing a communication primitive that uses a p…

cs.DC2019

Partitioned Paxos via the Network Data Plane

Huynh Tu Dang, Pietro Bressana, Han Wang +6

Consensus protocols are the foundation for building fault-tolerant, distributed systems, and services. They are also widely acknowledged as performance bottlenecks. Several recent…

cs.DC2016

Network Hardware-Accelerated Consensus

Huynh Tu Dang, Pietro Bressana, Han Wang +5

Consensus protocols are the foundation for building many fault-tolerant distributed systems and services. This paper posits that there are significant performance benefits to be ga…