6 papers
Structure-Aware Tensorial Model Reduction
Arjun Vijaywargiya, Eric C. Cyr, Anthony Gruber
This work investigates a two-stage method for constructing projection-based reduced-order models (ROMs) of parameterized partial differential equations (PDEs). Based on established…
Muon in Vision Transformers: Optimizer-Recipe Interactions and Gradient Spectra
Ben S. Southworth, Shuai Jiang, Daniel McBride +2
Muon is a recently developed matrix-aware optimizer that has shown strong results in transformer training, but its behavior in vision transformers (ViTs) is not yet well understood…
A Hybridizable Neural Time Integrator for Stable Autoregressive Forecasting
Brooks Kinch, Xiaozhe Hu, Yilong Huang +6
For autoregressive modeling of chaotic dynamical systems over long time horizons, the stability of both training and inference is a major challenge in building scientific foundatio…
Domain-Decomposed Graph Neural Network Surrogate Modeling for Ice Sheets
Adrienne M. Propp, Mauro Perego, Eric C. Cyr +5
Accurate yet efficient surrogate models are essential for large-scale simulations of partial differential equations (PDEs), particularly for uncertainty quantification (UQ) tasks t…
Deriving Transformer Architectures as Implicit Multinomial Regression
Jonas A. Actor, Anthony Gruber, Eric C. Cyr
While attention has been empirically shown to improve model performance, it lacks a rigorous mathematical justification. This short paper establishes a novel connection between att…
Leveraging KANs for Expedient Training of Multichannel MLPs via Preconditioning and Geometric Refinement
Jonas A. Actor, Graham Harper, Ben Southworth +1
Multilayer perceptrons (MLPs) are a workhorse machine learning architecture, used in a variety of modern deep learning frameworks. However, recently Kolmogorov-Arnold Networks (KAN…