Practical Federated Learning without a Server
arXiv:2503.05509 · doi:10.1145/3721146.3721938
Abstract
Federated Learning (FL) enables end-user devices to collaboratively train ML models without sharing raw data, thereby preserving data privacy. In FL, a central parameter server coordinates the learning process by iteratively aggregating the trained models received from clients. Yet, deploying a central server is not always feasible due to hardware unavailability, infrastructure constraints, or operational costs. We present Plexus, a fully decentralized FL system for large networks that operates without the drawbacks originating from having a central server. Plexus distributes the responsibilities of model aggregation and sampling among participating nodes while avoiding network-wide coordination. We evaluate Plexus using realistic traces for compute speed, pairwise latency and network capacity. Our experiments on three common learning tasks and with up to 1000 nodes empirically show that Plexus reduces time-to-accuracy by 1.4-1.6x, communication volume by 15.8-292x and training resources needed for convergence by 30.5-77.9x compared to conventional decentralized learning algorithms.
To appear in the proceedings of EuroMLSys'25
References in corpus (23)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Communication-Efficient Learning of Deep Networks from Decentralized Data
- Towards Federated Learning at Scale: System Design
- Can Decentralized Algorithms Outperform Centralized Algorithms? A Case Study for Decentralized Parallel Stochastic Gradient Descent
- SAFA: a Semi-Asynchronous Protocol for Fast Federated Learning with Low Overhead
- Gossip Learning with Linear Models on Fully Distributed Data
- Evaluating Gradient Inversion Attacks and Defenses in Federated Learning
- The Non-IID Data Quagmire of Decentralized Machine Learning
- Resource-Efficient Federated Learning
- Asynchronous Decentralized Parallel Stochastic Gradient Descent
- Oort: Efficient Federated Learning via Guided Participant Selection
- FedScale: Benchmarking Model and System Performance of Federated Learning at Scale
- Federated Learning with Buffered Asynchronous Aggregation
- Papaya: Practical, Private, and Scalable Federated Learning
- Decentralized Learning Made Easy with DecentralizePy
- Beyond spectral gap: The role of the topology in decentralized learning
- Decentralized Federated Learning: A Survey and Perspective
- Exponential Graph is Provably Efficient for Decentralized Deep Training
- Noiseless Privacy-Preserving Decentralized Learning
- Communication-Efficient Topologies for Decentralized Learning with Consensus Rate
- PeerSwap: A Peer-Sampler with Randomness Guarantees
- Beyond Exponential Graph: Communication-Efficient Topologies for Decentralized Learning via Finite-time Convergence
- Scalable Decentralized Learning with Teleportation