Communication-Efficient Edge AI: Algorithms and Systems
arXiv:2002.09668
Abstract
Artificial intelligence (AI) has achieved remarkable breakthroughs in a wide range of fields, ranging from speech processing, image classification to drug discovery. This is driven by the explosive growth of data, advances in machine learning (especially deep learning), and easy access to vastly powerful computing resources. Particularly, the wide scale deployment of edge devices (e.g., IoT devices) generates an unprecedented scale of data, which provides the opportunity to derive accurate models and develop various intelligent applications at the network edge. However, such enormous data cannot all be sent from end devices to the cloud for processing, due to the varying channel quality, traffic congestion and/or privacy concerns. By pushing inference and training processes of AI models to edge nodes, edge AI has emerged as a promising alternative. AI at the edge requires close cooperation among edge devices, such as smart phones and smart vehicles, and edge servers at the wireless access points and base stations, which however result in heavy communication overheads. In this paper, we present a comprehensive survey of the recent developments in various techniques for overcoming these communication challenges. Specifically, we first identify key communication challenges in edge AI systems. We then introduce communication-efficient techniques, from both algorithmic and system perspectives for training and inference tasks at the network edge. Potential future research directions are also highlighted.
This work has been submitted to the IEEE for possible publication
References in corpus (20)
- A Brief Survey of Deep Reinforcement Learning
- ZOO: Zeroth Order Optimization based Black-box Attacks to Deep Neural Networks without Training Substitute Models
- Deep Learning with Limited Numerical Precision
- Compressing Deep Convolutional Networks using Vector Quantization
- Compressing Neural Networks with the Hashing Trick
- Speeding up Convolutional Neural Networks with Low Rank Expansions
- Private federated learning on vertically partitioned data via entity resolution and additively homomorphic encryption
- Device Placement Optimization with Reinforcement Learning
- High-Dimensional Stochastic Gradient Quantization for Communication-Efficient Edge Learning
- AdaComp : Adaptive Residual Gradient Compression for Data-Parallel Distributed Training
- A Quasi-Newton Method Based Vertical Federated Learning Framework for Logistic Regression
- Reconfigurable-Intelligent-Surface Empowered Wireless Communications: Challenges and Opportunities
- Communication Efficient Federated Learning over Multiple Access Channels
- Gossip training for deep learning
- Iterative Hessian sketch: Fast and accurate solution approximation for constrained least-squares
- Reconfigurable Intelligent Surface for Green Edge Inference
- Distributed Sensing with Orthogonal Multiple Access: To code or not to Code?
- Distributed Stochastic Gradient Descent Using LDGM Codes
- Communication constrained cloud-based long-term visual localization in real time
- Distributed Deep Learning Strategies For Automatic Speech Recognition
Cited by in corpus (5)
- Wireless for Machine Learning
- Communication-Computation Trade-Off in Resource-Constrained Edge Inference
- Machine Learning for Predictive Deployment of UAVs with Multiple Access
- Branchy-GNN: a Device-Edge Co-Inference Framework for Efficient Point Cloud Processing
- Robust error bounds for quantised and pruned neural networks