ACDC: A Structured Efficient Linear Layer
arXiv:1511.05946
Abstract
The linear layer is one of the most pervasive modules in deep learning representations. However, it requires parameters and operations. These costs can be prohibitive in mobile applications or prevent scaling in many domains. Here, we introduce a deep, differentiable, fully-connected neural network module composed of diagonal matrices of parameters, and , and the discrete cosine transform . The core module, structured as , has parameters and incurs operations. We present theoretical results showing how deep cascades of ACDC layers approximate linear layers. ACDC is, however, a stand-alone module and can be used in combination with any other types of module. In our experiments, we show that it can indeed be successfully interleaved with ReLU modules in convolutional neural networks for image recognition. Our experiments also study critical factors in the training of these structured modules, including initialization and depth. Finally, this paper also provides a connection between structured linear transforms used in deep learning and the field of Fourier optics, illustrating how ACDC could in principle be implemented with lenses and diffractive elements.
References in corpus (10)
- Distilling the Knowledge in a Neural Network
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- Weight Uncertainty in Neural Networks
- Compressing Deep Convolutional Networks using Vector Quantization
- Compressing Neural Networks with the Hashing Trick
- Memory Bounded Deep Convolutional Networks
- Random Projections through multiple optical scattering: Approximating kernels at the speed of light
- Towards Trainable Media: Using Waves for Neural Network-Style Training
- Speeding Up Neural Networks for Large Scale Classification using WTA Hashing
Cited by in corpus (13)
- Multiplicative LSTM for sequence modelling
- Generalisation error in learning with random features and the hidden manifold model
- Reservoir Computing meets Recurrent Kernels and Structured Transforms
- Training Sparse Neural Networks
- GST: Group-Sparse Training for Accelerating Deep Reinforcement Learning
- Learning Low-rank Deep Neural Networks via Singular Vector Orthogonality Regularization and Singular Value Sparsification
- Understanding and Training Deep Diagonal Circulant Neural Networks
- Text Generation with Exemplar-based Adaptive Decoding
- Rethinking Neural Operations for Diverse Tasks
- Initialization and Regularization of Factorized Neural Layers
- CircConv: A Structured Convolution with Low Complexity
- Is the Number of Trainable Parameters All That Actually Matters?
- Training compact deep learning models for video classification using circulant matrices