HAT: Hardware-Aware Transformers for Efficient Natural Language Processing
arXiv:2005.14187 · doi:10.18653/v1/2020.acl-main.686
Abstract
Transformers are ubiquitous in Natural Language Processing (NLP) tasks, but they are difficult to be deployed on hardware due to the intensive computation. To enable low-latency inference on resource-constrained hardware platforms, we propose to design Hardware-Aware Transformers (HAT) with neural architecture search. We first construct a large design space with and . Then we train a that covers all candidates in the design space, and efficiently produces many with weight sharing. Finally, we perform an evolutionary search with a hardware latency constraint to find a specialized dedicated to run fast on the target hardware. Extensive experiments on four machine translation tasks demonstrate that HAT can discover efficient models for different hardware (CPU, GPU, IoT device). When running WMT'14 translation task on Raspberry Pi-4, HAT can achieve speedup, smaller size over baseline Transformer; speedup, smaller size over Evolved Transformer with less search cost and no performance loss. HAT code is https://github.com/mit-han-lab/hardware-aware-transformers.git
Accepted to ACL 2020. 14 pages, 12 figures. Code available at http://github.com/mit-han-lab/hardware-aware-transformers.git
References in corpus (6)
- Neural Architecture Search with Reinforcement Learning
- Pay Less Attention with Lightweight and Dynamic Convolutions
- GCN-RL Circuit Designer: Transferable Transistor Sizing with Graph Neural Networks and Reinforcement Learning
- SpArch: Efficient Architecture for Sparse Matrix Multiplication
- The Evolved Transformer
- Lite Transformer with Long-Short Range Attention
Cited by in corpus (37)
- Transformers in Vision: A Survey
- SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning
- QuantumNAS: Noise-Adaptive Search for Robust Quantum Circuits
- Enable Deep Learning on Mobile Devices: Methods, Systems, and Applications
- DynaBERT: Dynamic BERT with Adaptive Width and Depth
- A Comprehensive Survey on Hardware-Aware Neural Architecture Search
- PVNAS: 3D Neural Architecture Search with Point-Voxel Convolution
- NAS-BERT: Task-Agnostic and Adaptive-Size BERT Compression with Neural Architecture Search
- Searching Efficient 3D Architectures with Sparse Point-Voxel Convolution
- Mokey: Enabling Narrow Fixed-Point Inference for Out-of-the-Box Floating-Point Transformer Models
- One Proxy Device Is Enough for Hardware-Aware Neural Architecture Search
- STI: Turbocharge NLP Inference at the Edge via Elastic Pipelining
- Convolution-enhanced Evolving Attention Networks
- Pipeline Parallelism for Inference on Heterogeneous Edge Computing
- Searching for Efficient Multi-Stage Vision Transformers
- GOBO: Quantizing Attention-Based NLP Models for Low Latency and Energy Efficient Inference
- You Only Compress Once: Towards Effective and Elastic BERT Compression via Exploit-Explore Stochastic Nature Gradient
- Dynamic Multi-Branch Layers for On-Device Neural Machine Translation
- Dancing along Battery: Enabling Transformer with Run-time Reconfigurability on Mobile Devices
- Pay Attention when Required
- Dynamic-OFA: Runtime DNN Architecture Switching for Performance Scaling on Heterogeneous Embedded Platforms
- Hybrid SLC-MLC RRAM Mixed-Signal Processing-in-Memory Architecture for Transformer Acceleration via Gradient Redistribution
- Generic Neural Architecture Search via Regression
- TransAxx: Efficient Transformers with Approximate Computing
- Memory-Efficient Differentiable Transformer Architecture Search
- MS-Net: A Multi-Path Sparse Model for Motion Prediction in Multi-Scenes
- NASA: Neural Architecture Search and Acceleration for Hardware Inspired Hybrid Networks
- Scaling Up Deep Neural Network Optimization for Edge Inference
- Understanding and Overcoming the Challenges of Efficient Transformer Quantization
- AutoTinyBERT: Automatic Hyper-parameter Optimization for Efficient Pre-trained Language Models
- The NiuTrans System for the WMT21 Efficiency Task
- EfficientBERT: Progressively Searching Multilayer Perceptron via Warm-up Knowledge Distillation
- RankNAS: Efficient Neural Architecture Search by Pairwise Ranking
- Searching for TrioNet: Combining Convolution with Local and Global Self-Attention
- Enabling Design Methodologies and Future Trends for Edge AI: Specialization and Co-design
- MicroNet for Efficient Language Modeling
- LV-BERT: Exploiting Layer Variety for BERT