3 papers
math.OC2026
A Non-Monotone Preconditioned Trust-Region Method for Neural Network Training
Andrea Angino, Bindi Ãapriqi, Shega Likaj +2
Training deep neural networks at scale can benefit from domain decomposition, where the network is split into subdomains trained in parallel and coupled by a global trust-region me…
math.NA2026
Multi-Preconditioned LBFGS for Training Finite-Basis PINNs
Marc Salvadó-Benasco, Aymane Kssim, Alexander Heinlein +3
A multi-preconditioned LBFGS (MP-LBFGS) algorithm is introduced for training finite-basis physics-informed neural networks (FBPINNs). The algorithm is motivated by the nonlinear ad…
cs.LG2026
Layer-Parallel Training for Transformers
Shuai Jiang, Marc Salvadó-Benasco, Eric C. Cyr +3
We present a new training methodology for transformers using a multilevel, layer-parallel approach. Through a neural ODE formulation of transformers, our application of a multileve…