Expressiveness of Rectifier Networks
arXiv:1511.05678
Abstract
Rectified Linear Units (ReLUs) have been shown to ameliorate the vanishing gradient problem, allow for efficient backpropagation, and empirically promote sparsity in the learned parameters. They have led to state-of-the-art results in a variety of applications. However, unlike threshold and sigmoid networks, ReLU networks are less explored from the perspective of their expressiveness. This paper studies the expressiveness of ReLU networks. We characterize the decision boundary of two-layer ReLU networks by constructing functionally equivalent threshold networks. We show that while the decision boundary of a two-layer ReLU network can be captured by a threshold network, the latter may require an exponentially larger number of hidden units. We also formulate sufficient conditions for a corresponding logarithmic reduction in the number of hidden units to represent a sign network as a ReLU network. Finally, we experimentally compare threshold networks and their much smaller ReLU counterparts with respect to their ability to learn from synthetically generated data.
Published in ICML 2016. Supplementary material included
References in corpus (2)
Cited by in corpus (10)
- Knowledge Distillation in Deep Learning and its Applications
- Neural Policy Gradient Methods: Global Optimality and Rates of Convergence
- On the Expressive Power of Deep Neural Networks
- Knowledge Transfer via Distillation of Activation Boundaries Formed by Hidden Neurons
- A Comprehensive Overhaul of Feature Distillation
- Hierarchical Decomposition of Nonlinear Dynamics and Control for System Identification and Policy Distillation
- Enhancing Data-Free Adversarial Distillation with Activation Regularization and Virtual Interpolation
- Learning Constraints for Structured Prediction Using Rectifier Networks
- Learning Constraints and Descriptive Segmentation for Subevent Detection
- A Neural Frequency-Severity Model and Its Application to Insurance Claims