2 papers
stat.ML2026
Path-conditioned training: a principled way to rescale ReLU neural networks
Arthur Lebeurrier, Titouan Vayer, Rémi Gribonval
Despite recent algorithmic advances, we still lack principled ways to leverage the well-documented rescaling symmetries in ReLU neural network parameters. While two properly rescal…
cs.LG2025
A multilevel approach to accelerate the training of Transformers
Guillaume Lauga, Maël Chaumette, Edgar Desainte-Maréville +2
In this article, we investigate the potential of multilevel approaches to accelerate the training of transformer architectures. Using an ordinary differential equation (ODE) interp…