Split Computing and Early Exiting for Deep Learning Applications: Survey and Research Challenges
arXiv:2103.04505 · doi:10.1145/3527155
Abstract
Mobile devices such as smartphones and autonomous vehicles increasingly rely on deep neural networks (DNNs) to execute complex inference tasks such as image classification and speech recognition, among others. However, continuously executing the entire DNN on mobile devices can quickly deplete their battery. Although task offloading to cloud/edge servers may decrease the mobile device's computational burden, erratic patterns in channel quality, network, and edge server load can lead to a significant delay in task execution. Recently, approaches based on split computing (SC) have been proposed, where the DNN is split into a head and a tail model, executed respectively on the mobile device and on the edge server. Ultimately, this may reduce bandwidth usage as well as energy consumption. Another approach, called early exiting (EE), trains models to embed multiple "exits" earlier in the architecture, each providing increasingly higher target accuracy. Therefore, the trade-off between accuracy and delay can be tuned according to the current conditions or application demands. In this paper, we provide a comprehensive survey of the state of the art in SC and EE strategies by presenting a comparison of the most relevant approaches. We conclude the paper by providing a set of compelling research challenges.
Accepted to ACM Computing Surveys (CSUR)
References in corpus (15)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Distilling the Knowledge in a Neural Network
- Rethinking Atrous Convolution for Semantic Image Segmentation
- Natural Language Processing (almost) from Scratch
- Neural Architecture Search with Reinforcement Learning
- SemEval-2017 Task 1: Semantic Textual Similarity - Multilingual and Cross-lingual Focused Evaluation
- SPINN: Synergistic Progressive Inference of Neural Networks over Device and Cloud
- Supervised Compression for Resource-Constrained Edge Computing Systems
- BottleFit: Learning Compressed Representations in Deep Neural Networks for Effective and Efficient Split Computing
- FastBERT: a Self-distilling BERT with Adaptive Inference Time
- How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers
- Lightweight compression of neural network feature tensors for collaborative intelligence
- Zero Time Waste: Recycling Predictions in Early Exit Neural Networks
- Single-Training Collaborative Object Detectors Adaptive to Bandwidth and Computation
- Packet-Loss-Tolerant Split Inference for Delay-Sensitive Deep Learning in Lossy Wireless Networks
Cited by in corpus (19)
- Adaptive Inference through Early-Exit Networks: Design, Challenges and Directions
- Architectural Vision for Quantum Computing in the Edge-Cloud Continuum
- I-SPLIT: Deep Network Interpretability for Split Computing
- FrankenSplit: Efficient Neural Feature Compression with Shallow Variational Bottleneck Injection for Mobile Edge Computing
- FOOL: Addressing the Downlink Bottleneck in Satellite Computing with Neural Feature Compression
- Semantic Edge Computing and Semantic Communications in 6G Networks: A Unifying Survey and Research Challenges
- LimitNet: Progressive, Content-Aware Image Offloading for Extremely Weak Devices & Networks
- Sustainable Edge Intelligence Through Energy-Aware Early Exiting
- Conditional computation in neural networks: principles and research trends
- 3D Point Cloud Object Detection on Edge Devices for Split Computing
- NaviSlim: Adaptive Context-Aware Navigation and Sensing via Dynamic Slimmable Networks
- Slimmable Encoders for Flexible Split DNNs in Bandwidth and Resource Constrained IoT Systems
- Divide and Save: Splitting Workload Among Containers in an Edge Device to Save Energy and Time
- NaviSplit: Dynamic Multi-Branch Split DNNs for Efficient Distributed Autonomous Navigation
- Spatio-Temporal Split Learning for Autonomous Aerial Surveillance using Urban Air Mobility (UAM) Networks
- ParaDiS: Parallelly Distributable Slimmable Neural Networks
- ScissionLite: Accelerating Distributed Deep Neural Networks Using Transfer Layer
- Edge Intelligence in Civil Aviation: Paradigms, Techniques, and Applications
- Collaborative P4-SDN DDoS Detection and Mitigation with Early-Exit Neural Networks