9 papers
Vanilla SGD with Momentum Survives Heavy-Tailed Noise: Convergence Analysis without Gradient Clipping or Normalization
Ryusei Yamada, Naoki Sato, Hideaki Iiduka
Stochastic gradient descent (SGD) is a cornerstone of modern optimization. While its performance under heavy-tailed noise is often addressed through specialized modifications such…
Convergence Bound and Critical Batch Size of Muon Optimizer
Naoki Sato, Hiroki Naganuma, Hideaki Iiduka
Muon, a recently proposed optimizer that leverages the inherent matrix structure of neural network parameters, has demonstrated strong empirical performance, indicating its potenti…
Muon Converges under Heavy-Tailed Noise: Nonconvex Hölder-Smooth Empirical Risk Minimization
Hideaki Iiduka
Muon is a recently proposed optimizer that enforces orthogonality in parameter updates by projecting gradients onto the Stiefel manifold, leading to stable and efficient training i…
Improved Convergence Rates of Muon Optimizer for Nonconvex Optimization
Shuntaro Nagashima, Hideaki Iiduka
The Muon optimizer has recently attracted attention due to its orthogonalized first-order updates, and a deeper theoretical understanding of its convergence behavior is essential f…
Lipschitz Multiscale Deep Equilibrium Models: A Theoretically Guaranteed and Accelerated Approach
Naoki Sato, Hideaki Iiduka
Deep equilibrium models (DEQs) achieve infinitely deep network representations without stacking layers by exploring fixed points of layer transformations in neural networks. Such m…
Convergence Analysis of SGD under Expected Smoothness
Yuta Kawamoto, Hideaki Iiduka
Stochastic gradient descent (SGD) is the workhorse of large-scale learning, yet classical analyses rely on assumptions that can be either too strong (bounded variance) or too coars…