The Effects of Hyperparameters on SGD Training of Neural Networks
arXiv:1508.02788
Abstract
The performance of neural network classifiers is determined by a number of hyperparameters, including learning rate, batch size, and depth. A number of attempts have been made to explore these parameters in the literature, and at times, to develop methods for optimizing them. However, exploration of parameter spaces has often been limited. In this note, I report the results of large scale experiments exploring these different parameters and their interactions.
References in corpus (1)
Cited by in corpus (13)
- GSA-DenseNet121-COVID-19: a Hybrid Deep Learning Architecture for the Diagnosis of COVID-19 Disease based on Gravitational Search Optimization Algorithm
- Measuring the Effects of Data Parallelism on Neural Network Training
- Embedding Hard Physical Constraints in Neural Network Coarse-Graining of 3D Turbulence
- Exploiting the ConvLSTM: Human Action Recognition using Raw Depth Video-Based Recurrent Neural Networks
- Exploring the Design Space of Deep Convolutional Neural Networks at Large Scale
- Demystifying Learning Rate Policies for High Accuracy Training of Deep Neural Networks
- AnalogVNN: A fully modular framework for modeling and optimizing photonic neural networks
- A Comparative Study on Regularization Strategies for Embedding-based Neural Networks
- Study on the Large Batch Size Training of Neural Networks Based on the Second Order Gradient
- A Sensitivity Analysis of Attention-Gated Convolutional Neural Networks for Sentence Classification
- Operational Calculus for Differentiable Programming
- Exploring the Optimized Value of Each Hyperparameter in Various Gradient Descent Algorithms
- A Resizable Mini-batch Gradient Descent based on a Multi-Armed Bandit