A Full-Stack Search Technique for Domain Optimized Deep Learning Accelerators
arXiv:2105.12842 · doi:10.1145/3503222.3507767
Abstract
The rapidly-changing deep learning landscape presents a unique opportunity for building inference accelerators optimized for specific datacenter-scale workloads. We propose Full-stack Accelerator Search Technique (FAST), a hardware accelerator search framework that defines a broad optimization environment covering key design decisions within the hardware-software stack, including hardware datapath, software scheduling, and compiler passes such as operation fusion and tensor padding. In this paper, we analyze bottlenecks in state-of-the-art vision and natural language processing (NLP) models, including EfficientNet and BERT, and use FAST to design accelerators capable of addressing these bottlenecks. FAST-generated accelerators optimized for single workloads improve Perf/TDP by 3.7x on average across all benchmarks compared to TPU-v3. A FAST-generated accelerator optimized for serving a suite of workloads improves Perf/TDP by 2.4x on average compared to TPU-v3. Our return on investment analysis shows that FAST-generated accelerators can potentially be practical for moderate-sized datacenter deployments.
Fixed typo
References in corpus (9)
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
- Bayesian Optimization with Unknown Constraints
- Mind Mappings: Enabling Efficient Algorithm-Accelerator Mapping Space Search
- Analytical Characterization and Design Space Exploration for Optimization of CNNs
- A Learned Performance Model for Tensor Processing Units
- Rethinking Co-design of Neural Architectures and Hardware Accelerators
- Apollo: Transferable Architecture Exploration
- FPGA/DNN Co-Design: An Efficient Design Methodology for IoT Intelligence on the Edge
- Learned Hardware/Software Co-Design of Neural Accelerators
Cited by in corpus (5)
- SambaNova SN40L: Scaling the AI Memory Wall with Dataflow and Composition of Experts
- The Dawn of AI-Native EDA: Opportunities and Challenges of Large Circuit Models
- NicePIM: Design Space Exploration for Processing-In-Memory DNN Accelerators with 3D-Stacked-DRAM
- Shared Memory-contention-aware Concurrent DNN Execution for Diversely Heterogeneous System-on-Chips
- Special Session: Towards an Agile Design Methodology for Efficient, Reliable, and Secure ML Systems