Searching for Activation Functions
arXiv:1710.05941
Abstract
The choice of activation functions in deep networks has a significant effect on the training dynamics and task performance. Currently, the most successful and widely-used activation function is the Rectified Linear Unit (ReLU). Although various hand-designed alternatives to ReLU have been proposed, none have managed to replace it due to inconsistent gains. In this work, we propose to leverage automatic search techniques to discover new activation functions. Using a combination of exhaustive and reinforcement learning-based search, we discover multiple novel activation functions. We verify the effectiveness of the searches by conducting an empirical evaluation with the best discovered activation function. Our experiments show that the best discovered activation function, , which we name Swish, tends to work better than ReLU on deeper models across a number of challenging datasets. For example, simply replacing ReLUs with Swish units improves top-1 classification accuracy on ImageNet by 0.9\% for Mobile NASNet-A and 0.6\% for Inception-ResNet-v2. The simplicity of Swish and its similarity to ReLU make it easy for practitioners to replace ReLUs with Swish units in any neural network.
Updated version of "Swish: a Self-Gated Activation Function"
References in corpus (7)
- Neural Architecture Search with Reinforcement Learning
- Theoretical Models of Learning to Learn
- RL: Fast Reinforcement Learning via Slow Reinforcement Learning
- Layer Normalization
- Learning to reinforcement learn
- Learning Activation Functions to Improve Deep Neural Networks
- Neural Optimizer Search with Reinforcement Learning
Cited by in corpus (51)
- ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices
- CSPNet: A New Backbone that can Enhance Learning Capability of CNN
- Machine Learning Holography for 3D Particle Field Imaging
- Weight-Sharing Neural Architecture Search: A Battle to Shrink the Optimization Gap
- Context-aware Dynamics Model for Generalization in Model-Based Reinforcement Learning
- GLU Variants Improve Transformer
- Perpetual Motion: Generating Unbounded Human Motion
- HARK Side of Deep Learning -- From Grad Student Descent to Automated Machine Learning
- Soft-Root-Sign Activation Function
- Funnel Activation for Visual Recognition
- Physics-Constrained Bayesian Neural Network for Fluid Flow Reconstruction with Sparse and Noisy Data
- AReLU: Attention-based Rectified Linear Unit
- Expressive Priors in Bayesian Neural Networks: Kernel Combinations and Periodic Functions
- BETANAS: BalancEd TrAining and selective drop for Neural Architecture Search
- DVOLVER: Efficient Pareto-Optimal Neural Network Architecture Search
- Can weight sharing outperform random architecture search? An investigation with TuNAS
- An adaptive surrogate modeling based on deep neural networks for large-scale Bayesian inverse problems
- Complexity Measures for Neural Networks with General Activation Functions Using Path-based Norms
- Classification of Diabetic Retinopathy via Fundus Photography: Utilization of Deep Learning Approaches to Speed up Disease Detection
- Knowledge-based Radiation Treatment Planning: A Data-driven Method Survey
- S3NAS: Fast NPU-aware Neural Architecture Search Methodology
- A Note on Deepfake Detection with Low-Resources
- Learning to Generate Synthetic Data via Compositing
- Interpreting Neural Networks Using Flip Points
- Richer priors for infinitely wide multi-layer perceptrons
- Learning Compact Neural Networks Using Ordinary Differential Equations as Activation Functions
- Machine Learning Surrogates for Predicting Response of an Aero-Structural-Sloshing System
- Anomaly scores for generative models
- Input Hessian Regularization of Neural Networks
- Symmetrical Gaussian Error Linear Units (SGELUs)
- Quasi-Monte Carlo sampling for machine-learning partial differential equations
- TanhSoft -- a family of activation functions combining Tanh and Softplus
- MHVAE: a Human-Inspired Deep Hierarchical Generative Model for Multimodal Representation Learning
- Attention-Based Face AntiSpoofing of RGB Images, using a Minimal End-2-End Neural Network
- LiDAR ICPS-net: Indoor Camera Positioning based-on Generative Adversarial Network for RGB to Point-Cloud Translation
- Efficient Design of Neural Networks with Random Weights
- Regularized Flexible Activation Function Combinations for Deep Neural Networks
- AutoShrink: A Topology-aware NAS for Discovering Efficient Neural Architecture
- Multi-Activation Hidden Units for Neural Networks with Random Weights
- Supply-Power-Constrained Cable Capacity Maximization Using Deep Neural Networks
- Machine Learning Approach for Transforming Scattering Parameters to Complex Permittivity
- Multikernel activation functions: formulation and a case study
- NeuralScale: Efficient Scaling of Neurons for Resource-Constrained Deep Neural Networks
- Predicting Effective Diffusivity of Porous Media from Images by Deep Learning
- Learning and Inference in Imaginary Noise Models
- Unnormalized Variational Bayes
- A Wasserstein Minimum Velocity Approach to Learning Unnormalized Models
- RLCache: Automated Cache Management Using Reinforcement Learning
- Constraining Logits by Bounded Function for Adversarial Robustness
- JNR: Joint-based Neural Rig Representation for Compact 3D Face Modeling
- Speaker Representation Learning using Global Context Guided Channel and Time-Frequency Transformations