An exploration of parameter redundancy in deep networks with circulant projections
arXiv:1502.03436
Abstract
We explore the redundancy of parameters in deep neural networks by replacing the conventional linear projection in fully-connected layers with the circulant projection. The circulant structure substantially reduces memory footprint and enables the use of the Fast Fourier Transform to speed up the computation. Considering a fully-connected neural network layer with d input nodes, and d output nodes, this method improves the time complexity from O(d^2) to O(dlogd) and space complexity from O(d^2) to O(d). The space savings are particularly important for modern deep convolutional neural network architectures, where fully-connected layers typically contain more than 90% of the network parameters. We further show that the gradient computation and optimization of the circulant projections can be performed very efficiently. Our experiments on three standard datasets show that the proposed approach achieves this significant gain in storage and efficiency with minimal increase in error rate compared to neural networks with unstructured projections.
International Conference on Computer Vision (ICCV) 2015
References in corpus (10)
- Improving neural networks by preventing co-adaptation of feature detectors
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Training and Operation of an Integrated Neuromorphic Network Based on Metal-Oxide Memristors
- Going Deeper with Convolutions
- Speeding up Convolutional Neural Networks with Low Rank Expansions
- Memory Bounded Deep Convolutional Networks
- Fast, simple and accurate handwritten digit classification by training shallow neural network classifiers with the 'extreme learning machine' algorithm
- Compact Nonlinear Maps and Circulant Extensions
- Deep Fried Convnets
- Enhanced Image Classification With a Fast-Learning Shallow Convolutional Neural Network
Cited by in corpus (27)
- A Survey of Model Compression and Acceleration for Deep Neural Networks
- Compression of Deep Convolutional Neural Networks for Fast and Low Power Mobile Applications
- NISP: Pruning Networks using Neuron Importance Score Propagation
- Compact Bilinear Pooling
- NoScope: Optimizing Neural Network Queries over Video at Scale
- Deep Learning Convolutional Networks for Multiphoton Microscopy Vasculature Segmentation
- A Unified Framework of DNN Weight Pruning and Weight Clustering/Quantization Using ADMM
- BlockDrop: Dynamic Inference Paths in Residual Networks
- Compact Nonlinear Maps and Circulant Extensions
- Wide Compression: Tensor Ring Nets
- Communication-Efficient Edge AI: Algorithms and Systems
- Run-Time Efficient RNN Compression for Inference on Edge Devices
- NoiseOut: A Simple Way to Prune Neural Networks
- C-LSTM: Enabling Efficient LSTM using Structured Compression Techniques on FPGAs
- ACDC: A Structured Efficient Linear Layer
- Fine-Pruning: Joint Fine-Tuning and Compression of a Convolutional Network with Bayesian Optimization
- Training Sparse Neural Networks
- Fully-adaptive Feature Sharing in Multi-Task Networks with Applications in Person Attribute Classification
- Transformed Regularization for Learning Sparse Deep Neural Networks
- Sparse Neural Networks Topologies
- Multiscale Hierarchical Convolutional Networks
- Deep Learning Towards Mobile Applications
- Restructuring, Pruning, and Adjustment of Deep Models for Parallel Distributed Inference
- FFT-Based Deep Learning Deployment in Embedded Systems
- Balanced Quantization: An Effective and Efficient Approach to Quantized Neural Networks
- Efficient and Robust Machine Learning for Real-World Systems
- Faster Binary Embeddings for Preserving Euclidean Distances