Adaptive Federated Optimization
arXiv:2003.00295
Abstract
Federated learning is a distributed machine learning paradigm in which a large number of clients coordinate with a central server to learn a model without sharing their own training data. Standard federated optimization methods such as Federated Averaging (FedAvg) are often difficult to tune and exhibit unfavorable convergence behavior. In non-federated settings, adaptive optimization methods have had notable success in combating such issues. In this work, we propose federated versions of adaptive optimizers, including Adagrad, Adam, and Yogi, and analyze their convergence in the presence of heterogeneous data for general non-convex settings. Our results highlight the interplay between client heterogeneity and communication efficiency. We also perform extensive experiments on these methods and show that the use of adaptive optimizers can significantly improve the performance of federated learning.
Published as a conference paper at ICLR 2021
References in corpus (7)
- On the Convergence of Adam and Beyond
- One weird trick for parallelizing convolutional neural networks
- Towards Federated Learning at Scale: System Design
- How to Escape Saddle Points Efficiently
- Adaptive Gradient Methods with Dynamic Bound of Learning Rate
- Adaptive Bound Optimization for Online Convex Optimization
- A Generic Approach for Escaping Saddle points
Cited by in corpus (78)
- Tackling the Objective Inconsistency Problem in Heterogeneous Federated Optimization
- Ensemble Distillation for Robust Model Fusion in Federated Learning
- FedML: A Research Library and Benchmark for Federated Machine Learning
- FedBN: Federated Learning on Non-IID Features via Local Batch Normalization
- Personalized Federated Learning with Moreau Envelopes
- Inverting Gradients -- How easy is it to break privacy in federated learning?
- Group Knowledge Transfer: Federated Learning of Large CNNs at the Edge
- Federated Learning Based on Dynamic Regularization
- FedGraphNN: A Federated Learning System and Benchmark for Graph Neural Networks
- Mime: Mimicking Centralized Stochastic Algorithms in Federated Learning
- Federated Learning on Non-IID Data Silos: An Experimental Study
- Client Selection in Federated Learning: Convergence Analysis and Power-of-Choice Selection Strategies
- Distilled One-Shot Federated Learning
- Training Speech Recognition Models with Federated Learning: A Quality/Cost Framework
- Exploiting Shared Representations for Personalized Federated Learning
- FedCluster: Boosting the Convergence of Federated Learning via Cluster-Cycling
- A first look into the carbon footprint of federated learning
- FedCV: A Federated Learning Framework for Diverse Computer Vision Tasks
- FLASHE: Additively Symmetric Homomorphic Encryption for Cross-Silo Federated Learning
- FedScale: Benchmarking Model and System Performance of Federated Learning at Scale
- Federated Learning via Posterior Averaging: A New Perspective and Practical Algorithms
- Federated Learning with Buffered Asynchronous Aggregation
- The Skellam Mechanism for Differentially Private Federated Learning
- Papaya: Practical, Private, and Scalable Federated Learning
- SpreadGNN: Serverless Multi-task Federated Learning for Graph Neural Networks
- FedBE: Making Bayesian Model Ensemble Applicable to Federated Learning
- Federated Learning with Compression: Unified Analysis and Sharp Guarantees
- The Distributed Discrete Gaussian Mechanism for Federated Learning with Secure Aggregation
- Prototype Guided Federated Learning of Visual Feature Representations
- Federated Reconstruction: Partially Local Federated Learning
- FedCM: Federated Learning with Client-level Momentum
- Practical and Private (Deep) Learning without Sampling or Shuffling
- Local Adaptivity in Federated Learning: Convergence and Consistency
- FedJAX: Federated learning simulation with JAX
- On the Outsized Importance of Learning Rates in Local Update Methods
- FedEval: A Holistic Evaluation Framework for Federated Learning
- Improving Semi-supervised Federated Learning by Reducing the Gradient Diversity of Models
- Achieving Linear Speedup with Partial Worker Participation in Non-IID Federated Learning
- Towards Practical Adam: Non-Convexity, Convergence Theory, and Mini-Batch Acceleration
- Effective Federated Adaptive Gradient Methods with Non-IID Decentralized Data
- Secure Aggregation for Buffered Asynchronous Federated Learning
- SSFL: Tackling Label Deficiency in Federated Learning via Personalized Self-Supervision
- FedNL: Making Newton-Type Methods Applicable to Federated Learning
- FedSKETCH: Communication-Efficient and Private Federated Learning via Sketching
- Federated Composite Optimization
- Efficient and Private Federated Learning with Partially Trainable Networks
- Faster Non-Convex Federated Learning via Global and Local Momentum
- RingFed: Reducing Communication Costs in Federated Learning on Non-IID Data
- Federated Mixture of Experts
- Federated Transfer Learning with Dynamic Gradient Aggregation
- Federated Learning on Non-IID Data: A Survey
- Secure Byzantine-Robust Distributed Learning via Clustering
- Convergence and Accuracy Trade-Offs in Federated Learning and Meta-Learning
- Cost-Effective Federated Learning in Mobile Edge Networks
- Multi-VFL: A Vertical Federated Learning System for Multiple Data and Label Owners
- Anarchic Federated Learning
- IFedAvg: Interpretable Data-Interoperability for Federated Learning
- Federated Nonconvex Sparse Learning
- Policy-Based Federated Learning
- Towards Model Agnostic Federated Learning Using Knowledge Distillation
- Adaptive Serverless Learning
- Communication-Efficient Agnostic Federated Averaging
- Dubhe: Towards Data Unbiasedness with Homomorphic Encryption in Federated Learning Client Selection
- DP-REC: Private & Communication-Efficient Federated Learning
- Accelerating Federated Learning in Heterogeneous Data and Computational Environments
- EasyFL: A Low-code Federated Learning Platform For Dummies
- ABC-FL: Anomalous and Benign client Classification in Federated Learning
- An Operator Splitting View of Federated Learning
- Adaptive Differentially Private Empirical Risk Minimization
- Quantized Adam with Error Feedback
- CFedAvg: Achieving Efficient Communication and Fast Convergence in Non-IID Federated Learning
- CADA: Communication-Adaptive Distributed Adam
- An Expectation-Maximization Perspective on Federated Learning
- Federated Unbiased Learning to Rank
- Federated Learning with Sparsification-Amplified Privacy and Adaptive Optimization
- Design and Analysis of Uplink and Downlink Communications for Federated Learning
- Dynamic Gradient Aggregation for Federated Domain Adaptation
- Compositional federated learning: Applications in distributionally robust averaging and meta learning