Structured Transforms for Small-Footprint Deep Learning
arXiv:1510.01722
Abstract
We consider the task of building compact deep learning pipelines suitable for deployment on storage and power constrained mobile devices. We propose a unified framework to learn a broad family of structured parameter matrices that are characterized by the notion of low displacement rank. Our structured transforms admit fast function and gradient evaluation, and span a rich range of parameter sharing configurations whose statistical modeling capacity can be explicitly tuned along a continuum from structured to unstructured. Experimental results show that these transforms can significantly accelerate inference and forward/backward passes during training, and offer superior accuracy-compactness-speed tradeoffs in comparison to a number of existing techniques. In keyword spotting applications in mobile speech recognition, our methods are much more effective than standard linear low-rank bottleneck layers and nearly retain the performance of state of the art models, while providing more than 3.5-fold compression.
To appear in NIPS 2015; 9 pages
References in corpus (4)
Cited by in corpus (34)
- Knowledge Distillation: A Survey
- A Survey of Model Compression and Acceleration for Deep Neural Networks
- Dynamic Convolutions: Exploiting Spatial Sparsity for Faster Inference
- Big-Little Net: An Efficient Multi-Scale Feature Representation for Visual and Speech Recognition
- Low-Memory Neural Network Training: A Technical Report
- Compressing RNNs for IoT devices by 15-38x using Kronecker Products
- Recycling Randomness with Structure for Sublinear time Kernel Expansions
- Structured LISTA for Multidimensional Harmonic Retrieval
- Binary embeddings with structured hashed projections
- Run-Time Efficient RNN Compression for Inference on Edge Devices
- Estimation of the sample covariance matrix from compressive measurements
- SiPPing Neural Networks: Sensitivity-informed Provable Pruning of Neural Networks
- Centripetal SGD for Pruning Very Deep Convolutional Networks with Complicated Structure
- Intermittent Learning: On-Device Machine Learning on Intermittently Powered System
- Structured Convolution Matrices for Energy-efficient Deep learning
- Scatterbrain: Unifying Sparse and Low-rank Attention Approximation
- GST: Group-Sparse Training for Accelerating Deep Reinforcement Learning
- Kaleidoscope: An Efficient, Learnable Representation For All Structured Linear Maps
- Sparse Linear Networks with a Fixed Butterfly Structure: Theory and Practice
- Dependency Aware Filter Pruning
- Building Efficient Deep Neural Networks with Unitary Group Convolutions
- Data-Dependent Coresets for Compressing Neural Networks with Applications to Generalization Bounds
- Accelerate CNN via Recursive Bayesian Pruning
- Compressing Language Models using Doped Kronecker Products
- Manipulating Identical Filter Redundancy for Efficient Pruning on Deep and Complicated CNN
- Deep Networks with Fast Retraining
- Separable Layers Enable Structured Efficient Linear Substitutions
- Learning Compressed Transforms with Low Displacement Rank
- Reinforcement Learning with Chromatic Networks for Compact Architecture Search
- Doping: A technique for efficient compression of LSTM models using sparse structured additive matrices
- Robust Student Network Learning
- CircConv: A Structured Convolution with Low Complexity
- Depthwise Non-local Module for Fast Salient Object Detection Using a Single Thread
- Low-bit Quantization of Recurrent Neural Network Language Models Using Alternating Direction Methods of Multipliers