Deep Tensor Convolution on Multicores
arXiv:1611.06565
Abstract
Deep convolutional neural networks (ConvNets) of 3-dimensional kernels allow joint modeling of spatiotemporal features. These networks have improved performance of video and volumetric image analysis, but have been limited in size due to the low memory ceiling of GPU hardware. Existing CPU implementations overcome this constraint but are impractically slow. Here we extend and optimize the faster Winograd-class of convolutional algorithms to the -dimensional case and specifically for CPU hardware. First, we remove the need to manually hand-craft algorithms by exploiting the relaxed constraints and cheap sparse access of CPU memory. Second, we maximize CPU utilization and multicore scalability by transforming data matrices to be cache-aware, integer multiples of AVX vector widths. Treating 2-dimensional ConvNets as a special (and the least beneficial) case of our approach, we demonstrate a 5 to 25-fold improvement in throughput compared to previous state-of-the-art.
11 pages, 4 figures, 1 supplementary doc
References in corpus (11)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems
- Two-Stream Convolutional Networks for Action Recognition in Videos
- Deep Generative Image Models using a Laplacian Pyramid of Adversarial Networks
- cuDNN: Efficient Primitives for Deep Learning
- Indoor Semantic Segmentation using depth information
- Exponential expressivity in deep neural networks through transient chaos
- Fast Convolutional Nets With fbfft: A GPU Performance Evaluation
- Real-Time Single Image and Video Super-Resolution Using an Efficient Sub-Pixel Convolutional Neural Network
- Recursive Training of 2D-3D Convolutional Networks for Neuronal Boundary Detection
- ZNNi - Maximizing the Inference Throughput of 3D Convolutional Networks on Multi-Core CPUs and GPUs