Efficient Deep Learning: A Survey on Making Deep Learning Models Smaller, Faster, and Better
arXiv:2106.08962 · doi:10.1145/3578938
Abstract
Deep Learning has revolutionized the fields of computer vision, natural language understanding, speech recognition, information retrieval and more. However, with the progressive improvements in deep learning models, their number of parameters, latency, resources required to train, etc. have all have increased significantly. Consequently, it has become important to pay attention to these footprint metrics of a model as well, not just its quality. We present and motivate the problem of efficiency in deep learning, followed by a thorough survey of the five core areas of model efficiency (spanning modeling techniques, infrastructure, and hardware) and the seminal work there. We also present an experiment-based guide along with code, for practitioners to optimize their model training and deployment. We believe this is the first comprehensive survey in the efficient deep learning space that covers the landscape of model efficiency from modeling techniques to hardware support. Our hope is that this survey would provide the reader with the mental model and the necessary understanding of the field to apply generic efficiency techniques to immediately get significant improvements, and also equip them with ideas for further research and experimentation to achieve additional gains.
References in corpus (23)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Distilling the Knowledge in a Neural Network
- Sequence to Sequence Learning with Neural Networks
- Neural Architecture Search with Reinforcement Learning
- Language Models are Few-Shot Learners
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- PaLM: Scaling Language Modeling with Pathways
- Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding
- Paying More Attention to Attention: Improving the Performance of Convolutional Neural Networks via Attention Transfer
- Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations
- To prune, or not to prune: exploring the efficacy of pruning for model compression
- The State of Sparsity in Deep Neural Networks
- Scaling Autoregressive Models for Content-Rich Text-to-Image Generation
- Billion-scale semi-supervised learning for image classification
- Quasi-Recurrent Neural Networks
- Population Based Training of Neural Networks
- High-Performance Neural Networks for Visual Object Classification
- MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices
- ProjectionNet: Learning Efficient On-Device Deep Networks Using Neural Projections
- Distilling Large Language Models into Tiny and Effective Students using pQRNN
- Learning from a Teacher using Unlabeled Data
Cited by in corpus (33)
- Structured Pruning for Deep Convolutional Neural Networks: A survey
- A Comprehensive Overview and Comparative Analysis on Deep Learning Models: CNN, RNN, LSTM, GRU
- Model Pruning Enables Localized and Efficient Federated Learning for Yield Forecasting and Data Sharing
- AMMUS : A Survey of Transformer-based Pretrained Models in Natural Language Processing
- Recent Advances in Deep Learning for Channel Coding: A Survey
- Model Compression in Practice: Lessons Learned from Practitioners Creating On-device Machine Learning Experiences
- In Defense and Revival of Bayesian Filtering for Thermal Infrared Object Tracking
- ALGAN: Anomaly Detection by Generating Pseudo Anomalous Data via Latent Variables
- OnDev-LCT: On-Device Lightweight Convolutional Transformers towards federated learning
- Compress and Compare: Interactively Evaluating Efficiency and Behavior Across ML Model Compression Experiments
- Designing deep neural networks for driver intention recognition
- Talaria: Interactively Optimizing Machine Learning Models for Efficient Inference
- Full-Cycle Energy Consumption Benchmark for Low-Carbon Computer Vision
- EavesDroid: Eavesdropping User Behaviors via OS Side-Channels on Smartphones
- Toward Efficient Convolutional Neural Networks With Structured Ternary Patterns
- A Perspective on Deep Vision Performance with Standard Image and Video Codecs
- SpotV2Net: Multivariate Intraday Spot Volatility Forecasting via Vol-of-Vol-Informed Graph Attention Networks
- Energy Efficiency in AI for 5G and Beyond: A DeepRx Case Study
- Advances in Small-Footprint Keyword Spotting: A Comprehensive Review of Efficient Models and Algorithms
- Greenformer: Factorization Toolkit for Efficient Deep Neural Networks
- Green AI: A systematic review and meta-analysis of its definitions, lifecycle models, hardware and measurement attempts
- Solving continuum and rarefied flows using differentiable programming
- Beyond ImageNet: Understanding Cross-Dataset Robustness of Lightweight Vision Models
- Enabling Vibration-Based Gesture Recognition on Everyday Furniture via Energy-Efficient FPGA Implementation of 1D Convolutional Networks
- Effective Predictive Modeling for Emergency Department Visits and Evaluating Exogenous Variables Impact: Using Explainable Meta-learning Gradient Boosting
- NeuroLGP-SM: Scalable Surrogate-Assisted Neuroevolution for Deep Neural Networks
- Learn&Drop: Fast Learning of CNNs based on Layer Dropping
- Towards Sustainable Artificial Intelligence: A Comprehensive Review and Comparative Analysis of Deep Learning Models' Carbon Footprint
- Efficient CNN with uncorrelated Bag of Features pooling
- From Tiny Machine Learning to Tiny Deep Learning: A Survey
- TOGGLE: Temporal Logic-Guided Large Language Model Compression for Edge
- Energy-Aware Metaheuristics
- The Efficiency Misnomer