7 papers
Vanilla SGD with Momentum Survives Heavy-Tailed Noise: Convergence Analysis without Gradient Clipping or Normalization
Ryusei Yamada, Naoki Sato, Hideaki Iiduka
Stochastic gradient descent (SGD) is a cornerstone of modern optimization. While its performance under heavy-tailed noise is often addressed through specialized modifications such…
Convergence Bound and Critical Batch Size of Muon Optimizer
Naoki Sato, Hiroki Naganuma, Hideaki Iiduka
Muon, a recently proposed optimizer that leverages the inherent matrix structure of neural network parameters, has demonstrated strong empirical performance, indicating its potenti…
Lipschitz Multiscale Deep Equilibrium Models: A Theoretically Guaranteed and Accelerated Approach
Naoki Sato, Hideaki Iiduka
Deep equilibrium models (DEQs) achieve infinitely deep network representations without stacking layers by exploring fixed points of layer transformations in neural networks. Such m…
Using Stochastic Gradient Descent to Smooth Nonconvex Functions: Analysis of Implicit Graduated Optimization
Naoki Sato, Hideaki Iiduka
The graduated optimization approach is a method for finding global optimal solutions for nonconvex functions by using a function smoothing operation with stochastic noise. This pap…
Momentum Does Not Reduce Stochastic Noise in Stochastic Gradient Descent
Naoki Sato, Hideaki Iiduka
For nonconvex objective functions, including those found in training deep neural networks, stochastic gradient descent (SGD) with momentum is said to converge faster and have bette…
Explicit and Implicit Graduated Optimization in Deep Neural Networks
Naoki Sato, Hideaki Iiduka
Graduated optimization is a global optimization technique that is used to minimize a multimodal nonconvex function by smoothing the objective function with noise and gradually refi…