4 papers
On the Convergence Behavior of Preconditioned Gradient Descent Toward the Rich Learning Regime
Shuai Jiang, Alexey Voronin, Eric Cyr +1
Spectral bias, the tendency of neural networks to learn low frequencies first, can be both a blessing and a curse. While it enhances the generalization capabilities by suppressing…
Multilevel Training for Kolmogorov Arnold Networks
Ben S. Southworth, Jonas A. Actor, Graham Harper +1
Algorithmic speedup of training common neural architectures is made difficult by the lack of structure guaranteed by the function compositions inherent to such networks. In contras…
Layer-Parallel Training for Transformers
Shuai Jiang, Marc Salvadó-Benasco, Eric C. Cyr +3
We present a new training methodology for transformers using a multilevel, layer-parallel approach. Through a neural ODE formulation of transformers, our application of a multileve…
Parallel-in-Time Solution of Allen-Cahn Equations by Integrating Operator Learning into the Parareal Method
Yuwei Geng, Junqi Yin, Eric C. Cyr +2
While recent advances in deep learning have shown promising efficiency gains in solving time-dependent partial differential equations (PDEs), matching the accuracy of conventional…