6 papers
A Diffusion Approximation for Temporal-Difference Learning with Linear Features under Markovian Noise
M. Forzo, E. Monzio Compagnoni, A. Russo +1
Temporal difference (TD) learning with linear function approximation is a core method for policy evaluation. Its classical continuous-time description is an ordinary differential e…
Beyond a Single Explanation of the Adam--SGD Gap
Chenxiang Zhang, Rustem Islamov, Enea Monzio Compagnoni +3
Prior work has identified several factors that can contribute to the performance gap between Adam and SGD, spanning data aspects, architecture design, and optimization properties.…
On the Interaction of Batch Noise, Adaptivity, and Compression, under -Smoothness: An SDE Approach
Enea Monzio Compagnoni, Rustem Islamov, Frank Norbert Proske +3
Distributed stochastic optimization intertwines (i) stochastic gradient noise, (ii) communication compression, and (iii) adaptive/normalized updates. While each factor has been stu…
Adaptive Methods Are Preferable in High Privacy Settings: An SDE Perspective
Enea Monzio Compagnoni, Alessandro Stanghellini, Rustem Islamov +2
Differential Privacy (DP) is becoming central to large-scale training as privacy regulations tighten. We revisit how DP noise interacts with adaptivity in optimization through the…
Adaptive Methods through the Lens of SDEs: Theoretical Insights on the Role of Noise
Enea Monzio Compagnoni, Tianlin Liu, Rustem Islamov +3
Despite the vast empirical evidence supporting the efficacy of adaptive optimization methods in deep learning, their theoretical understanding is far from complete. This work intro…
Unbiased and Sign Compression in Distributed Learning: Comparing Noise Resilience via SDEs
Enea Monzio Compagnoni, Rustem Islamov, Frank Norbert Proske +1
Distributed methods are essential for handling machine learning pipelines comprising large-scale models and datasets. However, their benefits often come at the cost of increased co…