4 papers
Quantitative Gaussian-Process limits of Tensor Programs
Andrea Agazzi, Eloy Mosig GarcÃa, Dario Trevisan
We study the infinite-width Gaussian-process limit of random neural networks through the lens of tensor programs, and we provide a quantitative convergence theory in Wasserstein di…
Global Optimization via Softmin Energy Minimization
Andrea Agazzi, Vittorio Carlei, Marco Romito +1
Global optimization, particularly for non-convex functions with multiple local minima, poses significant challenges for traditional gradient-based methods. While metaheuristic appr…
Stochastic Scaling Limits and Synchronization by Noise in Deep Transformer Models
Andrea Agazzi, Giuseppe Bruno, Eloy Mosig GarcÃa +2
We prove pathwise convergence of the layerwise evolution of tokens in a finite-depth, finite-width transformer model with MultiLayer Perceptron (MLP) blocks to a continuous-time st…
Quantitative convergence of trained single layer neural networks to Gaussian processes
Eloy Mosig, Andrea Agazzi, Dario Trevisan
In this paper, we study the quantitative convergence of shallow neural networks trained via gradient descent to their associated Gaussian processes in the infinite-width limit. Whi…