Progressive Compressed Records: Taking a Byte out of Deep Learning Data
arXiv:1911.00472
Abstract
Deep learning accelerators efficiently train over vast and growing amounts of data, placing a newfound burden on commodity networks and storage devices. A common approach to conserve bandwidth involves resizing or compressing data prior to training. We introduce Progressive Compressed Records (PCRs), a data format that uses compression to reduce the overhead of fetching and transporting data, effectively reducing the training time required to achieve a target accuracy. PCRs deviate from previous storage formats by combining progressive compression with an efficient storage layout to view a single dataset at multiple fidelities---all without adding to the total dataset size. We implement PCRs and evaluate them on a range of datasets, training tasks, and hardware architectures. Our work shows that: (i) the amount of compression a dataset can tolerate exceeds 50% of the original encoding for many DL training tasks; (ii) it is possible to automatically and efficiently select appropriate compression levels for a given task; and (iii) PCRs enable tasks to readily access compressed data at runtime---utilizing as little as half the training bandwidth and thus potentially doubling training speed.
References in corpus (23)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- The NumPy array: a structure for efficient numerical computation
- A Tutorial on Principal Component Analysis
- Variational image compression with a scale hyperprior
- Deep Gradient Compression: Reducing the Communication Bandwidth for Distributed Training
- Mixed Precision Training
- YouTube-8M: A Large-Scale Video Classification Benchmark
- A Survey of Model Compression and Acceleration for Deep Neural Networks
- ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness
- Highly Scalable Deep Learning Training System with Mixed-Precision: Training ImageNet in Four Minutes
- A study of the effect of JPG compression on adversarial images
- The Convergence Rate of Neural Networks for Learned Functions of Different Frequencies
- Towards Image Understanding from Deep Compression without Decoding
- Image Classification at Supercomputer Scale
- Characterizing Deep-Learning I/O Workloads in TensorFlow
- TBD: Benchmarking and Analyzing Deep Neural Network Training
- Deep Neural Network Compression with Single and Multiple Level Quantization
- 3LC: Lightweight and Effective Traffic Compression for Distributed Machine Learning
- Scale MLPerf-0.6 models on Google TPU-v3 Pods
- Faster Neural Network Training with Data Echoing
- Band-limited Training and Inference for Convolutional Neural Networks
- Exploring the limits of Concurrency in ML Training on Google TPUs
- Discrepancy, Coresets, and Sketches in Machine Learning