The Convergence Behavior of Adam under Heavy-Tailed Noise
arXiv:2607.27383
The paper provides the first convergence guarantees for the standard vector-form Adam optimizer under heavy‑tailed stochastic noise, showing convergence to stationary points with suboptimal iteration complexity that improves when the domain radius is known.
Abstract
We establish the first convergence guarantees for the plain vector-form Adam optimizer under heavy-tailed stochastic noise. While several Adam variants are known to achieve optimal iteration complexity in bounded-variance nonsmooth nonconvex optimization, little is understood about their behavior when stochastic gradients admit only a bounded -th central moment for some , a setting increasingly observed in modern deep learning. To address this gap, we generalize the recent online-to-nonconvex conversion framework to accommodate heavy-tailed martingale-difference noise. Building on this generalized framework, we develop a discounted regret analysis for Adam, without restrictive parameter coupling. Our results show that Adam converges to -stationary points under heavy-tailed noise. However, it exhibits a suboptimal iteration complexity and -dependent convergence, a suboptimality that persists even in the bounded-variance case (). Specifically, the -dominant term in the iteration complexity for reaching in-expectation stationarity is for , which simplifies to when . When the domain radius is known and used to control the online-learner output, a standard setup in related literature, the convergence rate improves to match the optimal complexity. In this case, the -dominant iteration complexity is for , which simplifies to when . These findings provide new theoretical insight into the robustness and limitations of Adam in heavy-tailed regimes.
Accepted at the 42nd Conference on Uncertainty in Artificial Intelligence (UAI 2026)