4 papers
Lipschitz Multiscale Deep Equilibrium Models: A Theoretically Guaranteed and Accelerated Approach
Naoki Sato, Hideaki Iiduka
Deep equilibrium models (DEQs) achieve infinitely deep network representations without stacking layers by exploring fixed points of layer transformations in neural networks. Such m…
Convergence Analysis of SGD under Expected Smoothness
Yuta Kawamoto, Hideaki Iiduka
Stochastic gradient descent (SGD) is the workhorse of large-scale learning, yet classical analyses rely on assumptions that can be either too strong (bounded variance) or too coars…
Explicit and Implicit Graduated Optimization in Deep Neural Networks
Naoki Sato, Hideaki Iiduka
Graduated optimization is a global optimization technique that is used to minimize a multimodal nonconvex function by smoothing the objective function with noise and gradually refi…
Scaled Conjugate Gradient Method for Nonconvex Optimization in Deep Neural Networks
Naoki Sato, Koshiro Izumi, Hideaki Iiduka
A scaled conjugate gradient method that accelerates existing adaptive methods utilizing stochastic gradients is proposed for solving nonconvex optimization problems with deep neura…