Continuously Differentiable Exponential Linear Units
arXiv:1704.07483
Abstract
Exponential Linear Units (ELUs) are a useful rectifier for constructing deep learning architectures, as they may speed up and otherwise improve learning by virtue of not have vanishing gradients and by having mean activations near zero. However, the ELU activation as parametrized in [1] is not continuously differentiable with respect to its input when the shape parameter alpha is not equal to 1. We present an alternative parametrization which is C1 continuous for all values of alpha, making the rectifier easier to reason about and making alpha easier to tune. This alternative parametrization has several other useful properties that the original parametrization of ELU does not: 1) its derivative with respect to x is bounded, 2) it contains both the linear transfer function and ReLU as special cases, and 3) it is scale-similar with respect to alpha.
References in corpus (1)
Cited by in corpus (19)
- Smooth Adversarial Training
- SPACE: Unsupervised Object-Oriented Scene Representation via Spatial Attention and Decomposition
- Machine learning for molecular dynamics with strongly correlated electrons
- Padé Activation Units: End-to-end Learning of Flexible Activation Functions in Deep Networks
- Estimates on the generalization error of Physics Informed Neural Networks (PINNs) for approximating PDEs
- Machine learning nonequilibrium electron forces for adiabatic spin dynamics
- AReLU: Attention-based Rectified Linear Unit
- X-Linear Attention Networks for Image Captioning
- Do Neural Optimal Transport Solvers Work? A Continuous Wasserstein-2 Benchmark
- Improving Generative Imagination in Object-Centric World Models
- Generative Neurosymbolic Machines
- Smooth activations and reproducibility in deep networks
- Generated Loss, Augmented Training, and Multiscale VAE
- Longitudinal Deep Kernel Gaussian Process Regression
- HAN: Higher-order Attention Network for Spoken Language Understanding
- Machine learning dynamics of phase separation in correlated electron magnets
- Introducing the DOME Activation Functions
- KeyCLD: Learning Constrained Lagrangian Dynamics in Keypoint Coordinates from Images
- Interpretable Deep Learning for Stock Returns: A Consensus-Bottleneck Asset Pricing Model