Gossip Learning with Linear Models on Fully Distributed Data
arXiv:1109.1396 · doi:10.1002/cpe.2858
Abstract
Machine learning over fully distributed data poses an important problem in peer-to-peer (P2P) applications. In this model we have one data record at each network node, but without the possibility to move raw data due to privacy considerations. For example, user profiles, ratings, history, or sensor readings can represent this case. This problem is difficult, because there is no possibility to learn local models, the system model offers almost no guarantees for reliability, yet the communication cost needs to be kept low. Here we propose gossip learning, a generic approach that is based on multiple models taking random walks over the network in parallel, while applying an online learning algorithm to improve themselves, and getting combined via ensemble learning methods. We present an instantiation of this approach for the case of classification with linear models. Our main contribution is an ensemble learning method which---through the continuous combination of the models in the network---implements a virtual weighted voting mechanism over an exponential number of models at practically no extra cost as compared to independent random walks. We prove the convergence of the method theoretically, and perform extensive experiments on benchmark datasets. Our experimental analysis demonstrates the performance and robustness of the proposed approach.
The paper was published in the journal Concurrency and Computation: Practice and Experience http://onlinelibrary.wiley.com/journal/10.1002/%28ISSN%291532-0634 (DOI: http://dx.doi.org/10.1002/cpe.2858). The modifications are based on the suggestions from the reviewers
Cited by in corpus (19)
- A Survey on Distributed Machine Learning
- Topology-aware Federated Learning in Edge Computing: A Comprehensive Survey
- A Survey on Privacy for B5G/6G: New Privacy Challenges, and Research Directions
- Decentralized federated learning of deep neural networks on non-iid data
- Communication-Efficient Training Workload Balancing for Decentralized Multi-Agent Learning
- Implicit Model Specialization through DAG-based Decentralized Federated Learning
- Collaborative Distributed Machine Learning
- TEE-based decentralized recommender systems: The raw data sharing redemption
- Get More for Less in Decentralized Learning Systems
- Noiseless Privacy-Preserving Decentralized Learning
- On the Limit Performance of Floating Gossip
- Boosting Asynchronous Decentralized Learning with Model Fragmentation
- Accelerating MoE Model Inference with Expert Sharding
- Practical Federated Learning without a Server
- SecureGBM: Secure Multi-Party Gradient Boosting
- Fair Decentralized Learning
- De-DSI: Decentralised Differentiable Search Index
- Noise-Resilient Ensemble Learning using Evidence Accumulation Clustering
- Detection of Insider Attacks in Distributed Projected Subgradient Algorithms