Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
On the Convergence Behavior of Preconditioned Gradient Descent Toward the Rich Learning Regime
Shuai Jiang, Alexey Voronin, Eric Cyr +1
Spectral bias, the tendency of neural networks to learn low frequencies first, can be both a blessing and a curse. While it enhances the generalization capabilities by suppressing…
cs.LG2026
Multilevel Training for Kolmogorov Arnold Networks
Ben S. Southworth, Jonas A. Actor, Graham Harper +1
Algorithmic speedup of training common neural architectures is made difficult by the lack of structure guaranteed by the function compositions inherent to such networks. In contras…
cs.LG2026
Layer-Parallel Training for Transformers
Shuai Jiang, Marc Salvadó-Benasco, Eric C. Cyr +3
We present a new training methodology for transformers using a multilevel, layer-parallel approach. Through a neural ODE formulation of transformers, our application of a multileve…