Improving Neural Network Training using Dynamic Learning Rate Schedule for PINNs and Image Classification
arXiv:2507.21749 · doi:10.1016/j.mlwa.2025.100697
Abstract
Training neural networks can be challenging, especially as the complexity of the problem increases. Despite using wider or deeper networks, training them can be a tedious process, especially if a wrong choice of the hyperparameter is made. The learning rate is one of such crucial hyperparameters, which is usually kept static during the training process. Learning dynamics in complex systems often requires a more adaptive approach to the learning rate. This adaptability becomes crucial to effectively navigate varying gradients and optimize the learning process during the training process. In this paper, a dynamic learning rate scheduler (DLRS) algorithm is presented that adapts the learning rate based on the loss values calculated during the training process. Experiments are conducted on problems related to physics-informed neural networks (PINNs) and image classification using multilayer perceptrons and convolutional neural networks, respectively. The results demonstrate that the proposed DLRS accelerates training and improves stability.
10 pages
References in corpus (16)
- Adam: A Method for Stochastic Optimization
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding
- Artificial Neural Networks for Solving Ordinary and Partial Differential Equations
- MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems
- Mixed Precision Training
- A disciplined approach to neural network hyper-parameters: Part 1 -- learning rate, batch size, momentum, and weight decay
- Large Batch Training of Convolutional Networks
- Hyper-Parameter Optimization: A Review of Algorithms and Applications
- The State of Sparsity in Deep Neural Networks
- Thinking Fast and Slow with Deep Learning and Tree Search
- Root Mean Square Layer Normalization
- Understanding Gradient Clipping in Private SGD: A Geometric Perspective
- On the Computational Inefficiency of Large Batch Sizes for Stochastic Gradient Descent
- The exploding gradient problem demystified - definition, prevalence, impact, origin, tradeoffs, and solutions
- Neural network based approach for solving problems in plane wave duct acoustics