Learning Neural Network Classifiers with Low Model Complexity
arXiv:1707.09933
Abstract
Modern neural network architectures for large-scale learning tasks have substantially higher model complexities, which makes understanding, visualizing and training these architectures difficult. Recent contributions to deep learning techniques have focused on architectural modifications to improve parameter efficiency and performance. In this paper, we derive a continuous and differentiable error functional for a neural network that minimizes its empirical error as well as a measure of the model complexity. The latter measure is obtained by deriving a differentiable upper bound on the Vapnik-Chervonenkis (VC) dimension of the classifier layer of a class of deep networks. Using standard backpropagation, we realize a training rule that tries to minimize the error on training samples, while improving generalization by keeping the model complexity low. We demonstrate the effectiveness of our formulation (the Low Complexity Neural Network - LCNN) across several deep learning algorithms, and a variety of large benchmark datasets. We show that hidden layer neurons in the resultant networks learn features that are crisp, and in the case of image datasets, quantitatively sharper. Our proposed approach yields benefits across a wide range of architectures, in comparison to and in conjunction with methods such as Dropout and Batch Normalization, and our results strongly suggest that deep learning techniques can benefit from model complexity control methods such as the LCNN learning rule.
This work has been submitted to the IEEE for possible publication
References in corpus (9)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Understanding deep learning requires rethinking generalization
- Group Sparse Regularization for Deep Neural Networks
- Adding Gradient Noise Improves Learning for Very Deep Networks
- Net-Trim: Convex Pruning of Deep Neural Networks with Performance Guarantee
- NoiseOut: A Simple Way to Prune Neural Networks
- The Incredible Shrinking Neural Network: New Perspectives on Learning Representations Through The Lens of Pruning
- Learning a Fuzzy Hyperplane Fat Margin Classifier with Minimum VC dimension