Supervised Compression for Resource-Constrained Edge Computing Systems
arXiv:2108.11898 · doi:10.1109/WACV51458.2022.00100
Abstract
There has been much interest in deploying deep learning algorithms on low-powered devices, including smartphones, drones, and medical sensors. However, full-scale deep neural networks are often too resource-intensive in terms of energy and storage. As a result, the bulk part of the machine learning operation is therefore often carried out on an edge server, where the data is compressed and transmitted. However, compressing data (such as images) leads to transmitting information irrelevant to the supervised task. Another popular approach is to split the deep network between the device and the server while compressing intermediate features. To date, however, such split computing strategies have barely outperformed the aforementioned naive data compression baselines due to their inefficient approaches to feature compression. This paper adopts ideas from knowledge distillation and neural image compression to compress intermediate feature representations more efficiently. Our supervised compression approach uses a teacher model and a student model with a stochastic bottleneck and learnable prior for entropy coding (Entropic Student). We compare our approach to various neural image and feature compression baselines in three vision tasks and found that it achieves better supervised rate-distortion performance while maintaining smaller end-to-end latency. We furthermore show that the learned feature representations can be tuned to serve multiple downstream tasks.
Accepted to WACV 2022. Code and models are available at https://github.com/yoshitomo-matsubara/supervised-compression
References in corpus (8)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Distilling the Knowledge in a Neural Network
- Rethinking Atrous Convolution for Semantic Image Segmentation
- Learning Transferable Visual Models From Natural Language Supervision
- CompressAI: a PyTorch library and evaluation platform for end-to-end compression research
- torchdistill: A Modular, Configuration-Driven Framework for Knowledge Distillation
- Lossy Compression for Lossless Prediction
- Progressive Neural Image Compression with Nested Quantization and Latent Ordering
Cited by in corpus (13)
- Split Computing and Early Exiting for Deep Learning Applications: Survey and Research Challenges
- Edge Deep Learning in Computer Vision and Medical Diagnostics: A Comprehensive Survey
- BottleFit: Learning Compressed Representations in Deep Neural Networks for Effective and Efficient Split Computing
- I-SPLIT: Deep Network Interpretability for Split Computing
- FrankenSplit: Efficient Neural Feature Compression with Shallow Variational Bottleneck Injection for Mobile Edge Computing
- FOOL: Addressing the Downlink Bottleneck in Satellite Computing with Neural Feature Compression
- Semantic Edge Computing and Semantic Communications in 6G Networks: A Unifying Survey and Research Challenges
- LimitNet: Progressive, Content-Aware Image Offloading for Extremely Weak Devices & Networks
- Progressive Neural Compression for Adaptive Image Offloading under Timing Constraints
- Flexible Variable-Rate Image Feature Compression for Edge-Cloud Systems
- Slimmable Encoders for Flexible Split DNNs in Bandwidth and Resource Constrained IoT Systems
- torchdistill Meets Hugging Face Libraries for Reproducible, Coding-Free Deep Learning Studies: A Case Study on NLP
- Saliency Driven Imagery Preprocessing for Efficient Compression -- Industrial Paper