Learning Activation Functions to Improve Deep Neural Networks
arXiv:1412.6830
Abstract
Artificial neural networks typically have a fixed, non-linear activation function at each neuron. We have designed a novel form of piecewise linear activation function that is learned independently for each neuron using gradient descent. With this adaptive activation function, we are able to improve upon deep neural network architectures composed of static rectified linear units, achieving state-of-the-art performance on CIFAR-10 (7.51%), CIFAR-100 (30.83%), and a benchmark from high-energy physics involving Higgs boson decay modes.
Accepted as a workshop paper contribution at the International Conference on Learning Representations (ICLR) 2015
References in corpus (4)
Cited by in corpus (15)
- Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification
- Searching for Activation Functions
- Training Skinny Deep Neural Networks with Iterative Hard Thresholding Methods
- Doubly Convolutional Neural Networks
- Activation Ensembles for Deep Neural Networks
- Neural Networks with Smooth Adaptive Activation Functions for Regression
- A Probabilistic Framework for Nonlinearities in Stochastic Neural Networks
- Deep Learning applied to Road Traffic Speed forecasting
- CIFAR-10: KNN-based Ensemble of Classifiers
- Collaborative Layer-wise Discriminative Learning in Deep Neural Networks
- Gaussian Process Neurons Learn Stochastic Activation Functions
- Hardware-Driven Nonlinear Activation for Stochastic Computing Based Deep Convolutional Neural Networks
- Evolving Parsimonious Networks by Mixing Activation Functions
- Improving neural networks with bunches of neurons modeled by Kumaraswamy units: Preliminary study
- An Effective Training Method For Deep Convolutional Neural Network