Norm-preserving Orthogonal Permutation Linear Unit Activation Functions (OPLU)
arXiv:1604.02313
Abstract
We propose a novel activation function that implements piece-wise orthogonal non-linear mappings based on permutations. It is straightforward to implement, and very computationally efficient, also it has little memory requirements. We tested it on two toy problems for feedforward and recurrent networks, it shows similar performance to tanh and ReLU. OPLU activation function ensures norm preservance of the backpropagated gradients, therefore it is potentially good for the training of deep, extra deep, and recurrent neural networks.
Submitted to conference ICANN'2016
References in corpus (7)
- On the difficulty of training Recurrent Neural Networks
- Fast and Accurate Deep Network Learning by Exponential Linear Units (ELUs)
- Return of the Devil in the Details: Delving Deep into Convolutional Nets
- A Simple Way to Initialize Recurrent Networks of Rectified Linear Units
- Unitary Evolution Recurrent Neural Networks
- All you need is a good init
- Data-dependent Initializations of Convolutional Neural Networks
Cited by in corpus (8)
- On orthogonality and learning recurrent networks with long term dependencies
- Simple and Principled Uncertainty Estimation with Deterministic Deep Learning via Distance Awareness
- Sorting out Lipschitz function approximation
- Efficient Orthogonal Parametrisation of Recurrent Neural Networks Using Householder Reflections
- Some Theoretical Insights into Wasserstein GANs
- Shifting Mean Activation Towards Zero with Bipolar Activation Functions
- Universal Lipschitz Approximation in Bounded Depth Neural Networks
- Approximating Lipschitz continuous functions with GroupSort neural networks