AutoFreeze: Automatically Freezing Model Blocks to Accelerate Fine-tuning
arXiv:2102.01386
Abstract
With the rapid adoption of machine learning (ML), a number of domains now use the approach of fine tuning models which were pre-trained on a large corpus of data. However, our experiments show that even fine-tuning on models like BERT can take many hours even when using modern accelerators like GPUs. While prior work proposes limiting the number of layers that are fine-tuned, e.g., freezing all layers but the last layer, we find that such static approaches lead to reduced accuracy. We propose, AutoFreeze, a system that uses an adaptive approach to choose which layers are trained and show how this can accelerate model fine-tuning while preserving accuracy. We also develop mechanisms to enable efficient caching of intermediate activations which can reduce the forward computation time when performing fine-tuning. We extend AutoFreeze to perform distributed fine-tuning and design two execution modes that minimize cost and running time respectively. Our evaluation on ten NLP tasks shows that AutoFreeze, with caching enabled, can improve fine-tuning on a single GPU by up to 2.55x. On a 64 GPU cluster, for fine-tuning on the AG's news dataset, AutoFreeze is able to achieve up to 4.38x speedup when optimizing for end-to-end training time and 5.03x reduction in total cost when optimizing for efficiency, without affecting model accuracy.
References in corpus (13)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Reducing Transformer Depth on Demand with Structured Dropout
- Parameter-Efficient Transfer Learning for NLP
- PyTorch Distributed: Experiences on Accelerating Data Parallel Training
- BERT and PALs: Projected Attention Layers for Efficient Adaptation in Multi-Task Learning
- FastBERT: a Self-distilling BERT with Adaptive Inference Time
- What Would Elsa Do? Freezing Layers During Transformer Fine-Tuning
- Accelerating Training of Transformer-Based Language Models with Progressive Layer Dropping
- Model Compression with Multi-Task Knowledge Distillation for Web-scale Question Answering System
- Accordion: Adaptive Gradient Communication via Critical Learning Regime Identification
- Zero-Shot Transfer Learning for Event Extraction
- ODIN: Automated Drift Detection and Recovery in Video Analytics
- Do You Have the Right Scissors? Tailoring Pre-trained Language Models via Monte-Carlo Methods