Flare: Flexible In-Network Allreduce
arXiv:2106.15565 · doi:10.1145/3458817.3476178
Abstract
The allreduce operation is one of the most commonly used communication routines in distributed applications. To improve its bandwidth and to reduce network traffic, this operation can be accelerated by offloading it to network switches, that aggregate the data received from the hosts, and send them back the aggregated result. However, existing solutions provide limited customization opportunities and might provide suboptimal performance when dealing with custom operators and data types, with sparse data, or when reproducibility of the aggregation is a concern. To deal with these problems, in this work we design a flexible programmable switch by using as a building block PsPIN, a RISC-V architecture implementing the sPIN programming model. We then design, model, and analyze different algorithms for executing the aggregation on this architecture, showing performance improvements compared to state-of-the-art approaches.
References in corpus (4)
- Sparsity in Deep Learning: Pruning and growth for efficient inference and training in neural networks
- A Survey on Data Plane Programming with P4: Fundamentals, Advances, and Applied Research
- An Exhaustive Survey on P4 Programmable Data Plane Switches: Taxonomy, Applications, Challenges, and Future Trends
- NetReduce: RDMA-Compatible In-Network Reduction for Distributed DNN Training Acceleration
Cited by in corpus (5)
- ASTRA-sim2.0: Modeling Hierarchical Networks and Disaggregated Systems for Large-model Training at Scale
- Themis: A Network Bandwidth-Aware Collective Scheduling Policy for Distributed Training of DL Models
- Near-Optimal Wafer-Scale Reduce
- SDT: A Low-cost and Topology-reconfigurable Testbed for Network Research
- In-Network Collective Operations: Game Changer or Challenge for AI Workloads?