3 papers
math.OC2026
Training Infinitely Deep and Wide Transformers
Raphaël Barboni, Maarten V. de Hoop, Takashi Furuya +1
Transformers have become the dominant architecture in modern machine learning, yet the theoretical understanding of their training dynamics remains limited. This paper develops a r…
cs.LG2025
Ultra-fast feature learning for the training of two-layer neural networks in the two-timescale regime
Raphaël Barboni, Gabriel Peyré, François-Xavier Vialard
We study the convergence of gradient methods for the training of mean-field single-hidden-layer neural networks with square loss. For this high-dimensional and non-convex optimizat…
cs.LG2025
Understanding the training of infinitely deep and wide ResNets with Conditional Optimal Transport
Raphaël Barboni, Gabriel Peyré, François-Xavier Vialard
We study the convergence of gradient flow for the training of deep neural networks. If Residual Neural Networks are a popular example of very deep architectures, their training con…