Learning Combinations of Activation Functions
arXiv:1801.09403 · doi:10.1109/ICPR.2018.8545362
Abstract
In the last decade, an active area of research has been devoted to design novel activation functions that are able to help deep neural networks to converge, obtaining better performance. The training procedure of these architectures usually involves optimization of the weights of their layers only, while non-linearities are generally pre-specified and their (possible) parameters are usually considered as hyper-parameters to be tuned manually. In this paper, we introduce two approaches to automatically learn different combinations of base activation functions (such as the identity function, ReLU, and tanh) during the training phase. We present a thorough comparison of our novel approaches with well-known architectures (such as LeNet-5, AlexNet, and ResNet-56) on three standard datasets (Fashion-MNIST, CIFAR-10, and ILSVRC-2012), showing substantial improvements in the overall performance, such as an increase in the top-1 accuracy for AlexNet on ILSVRC-2012 of 3.01 percentage points.
6 pages, 3 figures. Published as a conference paper at ICPR 2018. Code: https://bitbucket.org/francux/learning_combinations_of_activation_functions
References in corpus (5)
- TensorFlow: A system for large-scale machine learning
- Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
- Achieving Human Parity in Conversational Speech Recognition
- Learning Activation Functions to Improve Deep Neural Networks
- Random Walk Initialization for Training Very Deep Feedforward Networks
Cited by in corpus (8)
- Learning Specialized Activation Functions for Physics-informed Neural Networks
- Padé Activation Units: End-to-end Learning of Flexible Activation Functions in Deep Networks
- AReLU: Attention-based Rectified Linear Unit
- ENN: A Neural Network with DCT Adaptive Activation Functions
- Neuronal diversity can improve machine learning for physics and beyond
- Effectiveness of Scaled Exponentially-Regularized Linear Units (SERLUs)
- Regularized Flexible Activation Function Combinations for Deep Neural Networks
- Initialization Using Perlin Noise for Training Networks with a Limited Amount of Data