5 papers
Dropout and Random Gradient Masking Are Asymptotically Equivalent in Large ResNets
Javier Maass, Lénaïc Chizat
Dropout and Random Gradient Masking (RaM) are two training techniques used to improve performance in deep learning. Both techniques inject randomness into the training dynamics, bu…
ResNets of All Shapes and Sizes: Convergence of Training Dynamics in the Large-scale Limit
Louis-Pierre Chaintron, Lénaïc Chizat, Javier Maass
We establish convergence of the training dynamics of residual neural networks (ResNets) to their joint infinite depth L, hidden width M, and embedding dimension D limit. Specifical…
The Hidden Width of Deep ResNets: Tight Error Bounds and Phase Diagram
Lénaïc Chizat
We study the gradient-based training of large-depth residual networks (ResNets) from standard random initializations. We show that infinite-depth ResNets behave as if they were inf…
Phase Diagram of Dropout for Two-Layer Neural Networks in the Mean-Field Regime
Lénaïc Chizat, Pierre Marion, Yerkin Yesbay
Dropout is a standard training technique for neural networks that consists of randomly deactivating units at each step of their gradient-based training. It is known to improve perf…
The Feature Speed Formula: a flexible approach to scale hyper-parameters of deep neural networks
Lénaïc Chizat, Praneeth Netrapalli
Deep learning succeeds by doing hierarchical feature learning, yet tuning hyper-parameters (HP) such as initialization scales, learning rates etc., only give indirect control over…