2 papers
stat.ML2026
Dropout and Random Gradient Masking Are Asymptotically Equivalent in Large ResNets
Javier Maass, Lénaïc Chizat
Dropout and Random Gradient Masking (RaM) are two training techniques used to improve performance in deep learning. Both techniques inject randomness into the training dynamics, bu…
stat.ML2026
ResNets of All Shapes and Sizes: Convergence of Training Dynamics in the Large-scale Limit
Louis-Pierre Chaintron, Lénaïc Chizat, Javier Maass
We establish convergence of the training dynamics of residual neural networks (ResNets) to their joint infinite depth L, hidden width M, and embedding dimension D limit. Specifical…