2 papers
math.PR2026
Stochastic Scaling Limits and Synchronization by Noise in Deep Transformer Models
Andrea Agazzi, Giuseppe Bruno, Eloy Mosig GarcÃa +2
We prove pathwise convergence of the layerwise evolution of tokens in a finite-depth, finite-width transformer model with MultiLayer Perceptron (MLP) blocks to a continuous-time st…
stat.ML2025
Effective continuous equations for adaptive SGD: a stochastic analysis view
Luca Callisti, Marco Romito, Francesco Triggiano
We present a theoretical analysis of some popular adaptive Stochastic Gradient Descent (SGD) methods in the small learning rate regime. Using the stochastic modified equations fram…