3 papers
cs.LG2025
Adaptive Batch Size and Learning Rate Scheduler for Stochastic Gradient Descent Based on Minimization of Stochastic First-order Oracle Complexity
Hikaru Umeda, Hideaki Iiduka
The convergence behavior of mini-batch stochastic gradient descent (SGD) is highly sensitive to the batch size and learning rate settings. Recent theoretical studies have identifie…
cs.LG2025
Optimal Growth Schedules for Batch Size and Learning Rate in SGD that Reduce SFO Complexity
Hikaru Umeda, Hideaki Iiduka
The unprecedented growth of deep learning models has enabled remarkable advances but introduced substantial computational bottlenecks. A key factor contributing to training efficie…
cs.LG2025
Increasing Both Batch Size and Learning Rate Accelerates Stochastic Gradient Descent
Hikaru Umeda, Hideaki Iiduka
The performance of mini-batch stochastic gradient descent (SGD) strongly depends on setting the batch size and learning rate to minimize the empirical loss in training the deep neu…