A Learned Performance Model for Tensor Processing Units
arXiv:2008.01040
Abstract
Accurate hardware performance models are critical to efficient code generation. They can be used by compilers to make heuristic decisions, by superoptimizers as a minimization objective, or by autotuners to find an optimal configuration for a specific program. However, they are difficult to develop because contemporary processors are complex, and the recent proliferation of deep learning accelerators has increased the development burden. We demonstrate a method of learning performance models from a corpus of tensor computation graph programs for Tensor Processing Unit (TPU) instances. We show that our learned model outperforms a heavily-optimized analytical performance model on two tasks -- tile-size selection and operator fusion -- and that it helps an autotuner discover faster programs in a setting where access to TPUs is limited or expensive.
A version will appear in the Proceedings of the 4th MLSys Conference, San Jose, CA, USA, 2021
References in corpus (11)
- Inductive Representation Learning on Large Graphs
- Graph Transformer Networks
- Learning to Optimize Tensor Programs
- Peephole: Predicting Network Performance Before Training
- A Learned Performance Model for Tensor Processing Units
- Reinforced Genetic Algorithm Learning for Optimizing Computation Graphs
- ProGraML: Graph-based Deep Learning for Program Optimization and Analysis
- Neural Predictor for Neural Architecture Search
- Chameleon: Adaptive Code Optimization for Expedited Deep Neural Network Compilation
- ReNAS:Relativistic Evaluation of Neural Architecture Search
- Predicting the Computational Cost of Deep Learning Models
Cited by in corpus (7)
- A Full-Stack Search Technique for Domain Optimized Deep Learning Accelerators
- A Learned Performance Model for Tensor Processing Units
- FusionStitching: Boosting Memory Intensive Computations for Deep Learning Workloads
- CDMPP: A Device-Model Agnostic Framework for Latency Prediction of Tensor Programs
- Optimizing DNN Compilation for Distributed Training with Joint OP and Tensor Fusion
- A Runtime-Based Computational Performance Predictor for Deep Neural Network Training
- Transfer Learning Across Heterogeneous Features For Efficient Tensor Program Generation