The Efficiency Misnomer
arXiv:2110.12894
Abstract
Model efficiency is a critical aspect of developing and deploying machine learning models. Inference time and latency directly affect the user experience, and some applications have hard requirements. In addition to inference costs, model training also have direct financial and environmental impacts. Although there are numerous well-established metrics (cost indicators) for measuring model efficiency, researchers and practitioners often assume that these metrics are correlated with each other and report only few of them. In this paper, we thoroughly discuss common cost indicators, their advantages and disadvantages, and how they can contradict each other. We demonstrate how incomplete reporting of cost indicators can lead to partial conclusions and a blurred or incomplete picture of the practical considerations of different models. We further present suggestions to improve reporting of efficiency metrics.
References in corpus (22)
- Scaling Laws for Neural Language Models
- Efficient Deep Learning: A Survey on Making Deep Learning Models Smaller, Faster, and Better
- Speeding up Convolutional Neural Networks with Low Rank Expansions
- The State of Sparsity in Deep Neural Networks
- Perceiver IO: A General Architecture for Structured Inputs & Outputs
- Long Range Arena: A Benchmark for Efficient Transformers
- Parameter-Efficient Transfer Learning for NLP
- Carbon Emissions and Large Neural Network Training
- The Cost of Training NLP Models: A Concise Overview
- Charformer: Fast Character Transformers via Gradient-based Subword Tokenization
- Compacter: Efficient Low-Rank Hypercomplex Adapter Layers
- Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers
- Multiscale Vision Transformers
- How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers
- Scalable and Efficient MoE Training for Multitask Multilingual Models
- ResNet strikes back: An improved training procedure in timm
- Scaling Vision with Sparse Mixture of Experts
- Scaling Laws for Transfer
- Primer: Searching for Efficient Transformers for Language Modeling
- OmniNet: Omnidirectional Representations from Transformers
- Efficient Nearest Neighbor Language Models
- SCENIC: A JAX Library for Computer Vision Research and Beyond