Deep Learning with S-shaped Rectified Linear Activation Units
arXiv:1512.07030
Abstract
Rectified linear activation units are important components for state-of-the-art deep convolutional networks. In this paper, we propose a novel S-shaped rectified linear activation unit (SReLU) to learn both convex and non-convex functions, imitating the multiple function forms given by the two fundamental laws, namely the Webner-Fechner law and the Stevens law, in psychophysics and neural sciences. Specifically, SReLU consists of three piecewise linear functions, which are formulated by four learnable parameters. The SReLU is learned jointly with the training of the whole deep network through back propagation. During the training phase, to initialize SReLU in different layers, we propose a "freezing" method to degenerate SReLU into a predefined leaky rectified linear unit in the initial several training epochs and then adaptively learn the good initial values. SReLU can be universally used in the existing deep networks with negligible additional parameters and computation cost. Experiments with two popular CNN architectures, Network in Network and GoogLeNet on scale-various benchmarks including CIFAR10, CIFAR100, MNIST and ImageNet demonstrate that SReLU achieves remarkable improvement compared to other activation functions.
Accepted by AAAI-16
References in corpus (4)
Cited by in corpus (24)
- A survey on modern trainable activation functions
- Short-term traffic flow forecasting with spatial-temporal correlation in a hybrid deep learning framework
- Efficient Continual Learning in Neural Networks with Embedding Regularization
- A simple and efficient architecture for trainable activation functions
- Padé Activation Units: End-to-end Learning of Flexible Activation Functions in Deep Networks
- ProbAct: A Probabilistic Activation Function for Deep Neural Networks
- Keep the Gradients Flowing: Using Gradient Flow to Study Sparse Network Optimization
- Soft-Root-Sign Activation Function
- Language Independent Single Document Image Super-Resolution using CNN for improved recognition
- Smooth activations and reproducibility in deep networks
- Learning Over Long Time Lags
- Activation function impact on Sparse Neural Networks
- TanhExp: A Smooth Activation Function with High Convergence Speed for Lightweight Neural Networks
- MATGANIP: Learning to Discover the Structure-Property Relationship in Perovskites with Generative Adversarial Networks
- Topological Insights into Sparse Neural Networks
- Deep ensembles based on Stochastic Activation Selection for Polyp Segmentation
- DeepLABNet: End-to-end Learning of Deep Radial Basis Networks with Fully Learnable Basis Functions
- Comparisons among different stochastic selection of activation layers for convolutional neural networks for healthcare
- Deep Global-Connected Net With The Generalized Multi-Piecewise ReLU Activation in Deep Learning
- Multikernel activation functions: formulation and a case study
- Overcoming Overfitting and Large Weight Update Problem in Linear Rectifiers: Thresholded Exponential Rectified Linear Units
- Shakeout: A New Approach to Regularized Deep Neural Network Training
- An Adaptive Deep Learning Algorithm Based Autoencoder for Interference Channels
- Generation and Simulation of Yeast Microscopy Imagery with Deep Learning