Fast Algorithms for Convolutional Neural Networks
arXiv:1509.09308
Abstract
Deep convolutional neural networks take GPU days of compute time to train on large data sets. Pedestrian detection for self driving cars requires very low latency. Image recognition for mobile phones is constrained by limited processing resources. The success of convolutional neural networks in these situations is limited by how fast we can compute them. Conventional FFT based convolution is fast for large filters, but state of the art convolutional neural networks use small, 3x3 filters. We introduce a new class of fast algorithms for convolutional neural networks using Winograd's minimal filtering algorithms. The algorithms compute minimal complexity convolution over small tiles, which makes them fast with small filters and small batch sizes. We benchmark a GPU implementation of our algorithm with the VGG network and show state of the art throughput at batch sizes from 1 to 64.
References in corpus (6)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Deep Learning with Limited Numerical Precision
- One weird trick for parallelizing convolutional neural networks
- Fast Training of Convolutional Networks through FFTs
- Powers of Tensors and Fast Matrix Multiplication
- maxDNN: An Efficient Convolution Kernel for Deep Learning with Maxwell GPUs
Cited by in corpus (35)
- TensorFlow: A system for large-scale machine learning
- Pruning Convolutional Neural Networks for Resource Efficient Inference
- A Survey of Model Compression and Acceleration for Deep Neural Networks
- Channel Pruning for Accelerating Very Deep Neural Networks
- Tensor Comprehensions: Framework-Agnostic High-Performance Machine Learning Abstractions
- Neural GPUs Learn Algorithms
- PVANET: Deep but Lightweight Neural Networks for Real-time Object Detection
- TVM: An Automated End-to-End Optimizing Compiler for Deep Learning
- Faster CNNs with Direct Sparse Convolutions and Guided Pruning
- Optimizing Performance of Recurrent Neural Networks on GPUs
- BranchyNet: Fast Inference via Early Exiting from Deep Neural Networks
- Structured Probabilistic Pruning for Convolutional Neural Network Acceleration
- ShiftCNN: Generalized Low-Precision Architecture for Inference of Convolutional Neural Networks
- MEC: Memory-efficient Convolution for Deep Neural Network
- Efficient Sparse-Winograd Convolutional Neural Networks
- Photonic Convolution Neural Network Based on Interleaved Time-Wavelength Modulation
- Exploration of Low Numeric Precision Deep Learning Inference Using Intel FPGAs
- Warped Convolutions: Efficient Invariance to Spatial Transformations
- Deep Learning at 15PF: Supervised and Semi-Supervised Classification for Scientific Data
- An OpenCL(TM) Deep Learning Accelerator on Arria 10
- Can Active Memory Replace Attention?
- Faster Asynchronous SGD
- Coordinating Filters for Faster Deep Neural Networks
- cltorch: a Hardware-Agnostic Backend for the Torch Deep Neural Network Library, Based on OpenCL
- A flexible FPGA accelerator for convolutional neural networks
- A Metaprogramming and Autotuning Framework for Deploying Deep Learning Applications
- WinoCNN: Kernel Sharing Winograd Systolic Array for Efficient Convolutional Neural Network Acceleration on FPGAs
- cuConv: A CUDA Implementation of Convolution for CNN Inference
- The Scalability for Parallel Machine Learning Training Algorithm: Dataset Matters
- Deep Neural Network Approximation using Tensor Sketching
- RRNet: Repetition-Reduction Network for Energy Efficient Decoder of Depth Estimation
- Applying the Roofline model for Deep Learning performance optimizations
- Training CNNs faster with Dynamic Input and Kernel Downsampling
- Faster Convolution Inference Through Using Pre-Calculated Lookup Tables
- Towards Design Space Exploration and Optimization of Fast Algorithms for Convolutional Neural Networks (CNNs) on FPGAs