MLPerf Training Benchmark
arXiv:1910.01500
Abstract
Machine learning (ML) needs industry-standard performance benchmarks to support design and competitive evaluation of the many emerging software and hardware solutions for ML. But ML training presents three unique benchmarking challenges absent from other domains: optimizations that improve training throughput can increase the time to solution, training is stochastic and time to solution exhibits high variance, and software and hardware systems are so diverse that fair benchmarking with the same binary, code, and even hyperparameters is difficult. We therefore present MLPerf, an ML benchmark that overcomes these challenges. Our analysis quantitatively evaluates MLPerf's efficacy at driving performance and scalability improvements across two rounds of results from multiple vendors.
MLSys 2020
References in corpus (9)
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- cuDNN: Efficient Primitives for Deep Learning
- One weird trick for parallelizing convolutional neural networks
- Trained Ternary Quantization
- The Loss Surfaces of Multilayer Networks
- Large Batch Training of Convolutional Networks
- Deep Learning Recommendation Model for Personalization and Recommendation Systems
- Fathom: Reference Workloads for Modern Deep Learning Methods
- Scalable Realistic Recommendation Datasets through Fractal Expansions
Cited by in corpus (8)
- DeepRecSys: A System for Optimizing End-To-End At-scale Neural Recommendation Inference
- HULK: An Energy Efficiency Benchmark Platform for Responsible Natural Language Processing
- BenchCouncil's View on Benchmarking AI and Other Emerging Workloads
- DLSpec: A Deep Learning Task Exchange Specification
- The Pitfall of Evaluating Performance on Emerging AI Accelerators
- Neural Model-based Optimization with Right-Censored Observations
- AIBench: An Agile Domain-specific Benchmarking Methodology and an AI Benchmark Suite
- Bosch Deep Learning Hardware Benchmark