Compression of Deep Neural Networks on the Fly
arXiv:1509.08745
Abstract
Thanks to their state-of-the-art performance, deep neural networks are increasingly used for object recognition. To achieve these results, they use millions of parameters to be trained. However, when targeting embedded applications the size of these models becomes problematic. As a consequence, their usage on smartphones or other resource limited devices is prohibited. In this paper we introduce a novel compression method for deep neural networks that is performed during the learning phase. It consists in adding an extra regularization term to the cost function of fully-connected layers. We combine this method with Product Quantization (PQ) of the trained weights for higher savings in storage consumption. We evaluate our method on two data sets (MNIST and CIFAR10), on which we achieve significantly larger compression rates than state-of-the-art methods.
References in corpus (10)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Exploiting Linear Structure Within Convolutional Networks for Efficient Evaluation
- Going Deeper with Convolutions
- Theano: new features and speed improvements
- Compressing Deep Convolutional Networks using Vector Quantization
- Deep Neural Networks Rival the Representation of Primate IT Cortex for Core Visual Object Recognition
- Stochastic Pooling for Regularization of Deep Convolutional Neural Networks
- Compressing Neural Networks with the Hashing Trick
- Speeding up Convolutional Neural Networks with Low Rank Expansions
- Compressing Convolutional Neural Networks