Distributed Deep Convolutional Neural Networks for the Internet-of-Things
arXiv:1908.01656 · doi:10.1109/TC.2021.3062227
Abstract
Severe constraints on memory and computation characterizing the Internet-of-Things (IoT) units may prevent the execution of Deep Learning (DL)-based solutions, which typically demand large memory and high processing load. In order to support a real-time execution of the considered DL model at the IoT unit level, DL solutions must be designed having in mind constraints on memory and processing capability exposed by the chosen IoT technology. In this paper, we introduce a design methodology aiming at allocating the execution of Convolutional Neural Networks (CNNs) on a distributed IoT application. Such a methodology is formalized as an optimization problem where the latency between the data-gathering phase and the subsequent decision-making one is minimized, within the given constraints on memory and processing load at the units level. The methodology supports multiple sources of data as well as multiple CNNs in execution on the same IoT system allowing the design of CNN-based applications demanding autonomy, low decision-latency, and high Quality-of-Service.
References in corpus (1)
Cited by in corpus (6)
- A Survey on Collaborative DNN Inference for Edge Intelligence
- Distributed CNN Inference on Resource-Constrained UAVs for Surveillance Systems: Design and Optimization
- RL-DistPrivacy: Privacy-Aware Distributed Deep Inference for low latency IoT systems
- Semantic Edge Computing and Semantic Communications in 6G Networks: A Unifying Survey and Research Challenges
- OSCAR-P and aMLLibrary: Profiling and Predicting the Performance of FaaS-based Applications in Computing Continua
- DCentNet: Decentralized Multistage Biomedical Signal Classification using Early Exits