A Survey on Green Deep Learning
arXiv:2111.05193
Abstract
In recent years, larger and deeper models are springing up and continuously pushing state-of-the-art (SOTA) results across various fields like natural language processing (NLP) and computer vision (CV). However, despite promising results, it needs to be noted that the computations required by SOTA models have been increased at an exponential rate. Massive computations not only have a surprisingly large carbon footprint but also have negative effects on research inclusiveness and deployment on real-world applications. Green deep learning is an increasingly hot research field that appeals to researchers to pay attention to energy usage and carbon emission during model training and inference. The target is to yield novel results with lightweight and efficient technologies. Many technologies can be used to achieve this goal, like model compression and knowledge distillation. This paper focuses on presenting a systematic review of the development of Green deep learning technologies. We classify these approaches into four categories: (1) compact networks, (2) energy-efficient training strategies, (3) energy-efficient inference approaches, and (4) efficient data usage. For each category, we discuss the progress that has been achieved and the unresolved challenges.
References in corpus (75)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Distilling the Knowledge in a Neural Network
- Natural Language Processing (almost) from Scratch
- How transferable are features in deep neural networks?
- Bootstrap your own latent: A new approach to self-supervised Learning
- Improved Baselines with Momentum Contrastive Learning
- Unsupervised Learning of Visual Features by Contrasting Cluster Assignments
- Paying More Attention to Attention: Improving the Performance of Convolutional Neural Networks via Attention Transfer
- Pre-trained Models for Natural Language Processing: A Survey
- Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations
- Compressing Deep Convolutional Networks using Vector Quantization
- ERNIE: Enhanced Representation through Knowledge Integration
- DenseNet: Implementing Efficient ConvNet Descriptor Pyramids
- Multilingual Denoising Pre-training for Neural Machine Translation
- TernausNet: U-Net with VGG11 Encoder Pre-Trained on ImageNet for Image Segmentation
- Layer Normalization
- Generating Long Sequences with Sparse Transformers
- Beyond English-Centric Multilingual Machine Translation
- Hyper-Parameter Optimization: A Review of Algorithms and Applications
- Massively Multitask Networks for Drug Discovery
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
- Like What You Like: Knowledge Distill via Neuron Selectivity Transfer
- Reformer: The Efficient Transformer
- Reducing Transformer Depth on Demand with Structured Dropout
- Toward Multilingual Neural Machine Translation with Universal Encoder and Decoder
- Deep Equilibrium Models
- CLEAR: Contrastive Learning for Sentence Representation
- Transfer Learning for Sequence Tagging with Hierarchical Recurrent Networks
- The Evolved Transformer
- The Reversible Residual Network: Backpropagation Without Storing Activations
- Understanding and Improving Layer Normalization
- Empower Sequence Labeling with Task-Aware Neural Language Model
- ImageNet pre-trained models with batch normalization
- Quantifying the Carbon Emissions of Machine Learning
- Multilingual Neural Machine Translation with Knowledge Distillation
- Pre-Trained Image Processing Transformer
- Rethinking Attention with Performers
- StructBERT: Incorporating Language Structures into Pre-training for Deep Language Understanding
- Block-Sparse Recurrent Neural Networks
- Memory-Efficient Implementation of DenseNets
- Feature-map-level Online Adversarial Knowledge Distillation
- Finetuned Language Models Are Zero-Shot Learners
- Zero-Cost Proxies for Lightweight NAS
- Effective Quantization Methods for Recurrent Neural Networks
- Efficient 8-Bit Quantization of Transformer Neural Machine Language Translation Model
- Pruning and Quantization for Deep Neural Network Acceleration: A Survey
- Recurrent Neural Networks With Limited Numerical Precision
- M6: A Chinese Multimodal Pretrainer
- BinaryBERT: Pushing the Limit of BERT Quantization
- Learning Student-Friendly Teacher Networks for Knowledge Distillation
- Pre-training Multilingual Neural Machine Translation by Leveraging Alignment Information
- Neural Architecture Search without Training
- Knowledge Adaptation: Teaching to Adapt
- Dynamic Neural Networks: A Survey
- Noisy Differentiable Architecture Search
- Variable Computation in Recurrent Neural Networks
- Learning Accurate Low-Bit Deep Neural Networks with Stochastic Quantization
- Learning Implicitly Recurrent CNNs Through Parameter Sharing
- Lessons on Parameter Sharing across Layers in Transformers
- FEED: Feature-level Ensemble for Knowledge Distillation
- Parameter-Efficient Transfer Learning with Diff Pruning
- Knowledge Flow: Improve Upon Your Teachers
- Learned Low Precision Graph Neural Networks
- Knowledge Projection for Deep Neural Networks
- Layer-wise training of deep networks using kernel similarity
- Early Exiting with Ensemble Internal Classifiers
- CAPT: Contrastive Pre-Training for Learning Denoised Sequence Representations
- Compressing Neural Language Models by Sparse Word Representations
- Does Knowledge Distillation Really Work?
- TernaryBERT: Distillation-aware Ultra-low Bit BERT
- Progressively Stacking 2.0: A Multi-stage Layerwise Training Method for BERT Training Speedup
- RomeBERT: Robust Training of Multi-Exit BERT
- KNAS: Green Neural Architecture Search
- The OoO VLIW JIT Compiler for GPU Inference
- Neural Parameter Allocation Search