Introducing the DOME Activation Functions
arXiv:2109.14798
Abstract
In this paper, we introduce a novel non-linear activation function that spontaneously induces class-compactness and regularization in the embedding space of neural networks. The function is dubbed DOME for Difference Of Mirrored Exponential terms. The basic form of the function can replace the sigmoid or the hyperbolic tangent functions as an output activation function for binary classification problems. The function can also be extended to the case of multi-class classification, and used as an alternative to the standard softmax function. It can also be further generalized to take more flexible shapes suitable for intermediate layers of a network. We empirically demonstrate the properties of the function. We also show that models using the function exhibit extra robustness against adversarial attacks.
16 pages, 9 figures
References in corpus (6)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Distilling the Knowledge in a Neural Network
- Show and Tell: Lessons learned from the 2015 MSCOCO Image Captioning Challenge
- Theoretically Principled Trade-off between Robustness and Accuracy
- Searching for Activation Functions
- Improving Adversarial Robustness via Promoting Ensemble Diversity