collaborators

6 papers

math.NA2026

Structure-Aware Tensorial Model Reduction

Arjun Vijaywargiya, Eric C. Cyr, Anthony Gruber

This work investigates a two-stage method for constructing projection-based reduced-order models (ROMs) of parameterized partial differential equations (PDEs). Based on established…

cs.LG2026

Muon in Vision Transformers: Optimizer-Recipe Interactions and Gradient Spectra

Ben S. Southworth, Shuai Jiang, Daniel McBride +2

Muon is a recently developed matrix-aware optimizer that has shown strong results in transformer training, but its behavior in vision transformers (ViTs) is not yet well understood…

cs.LG2026

A Hybridizable Neural Time Integrator for Stable Autoregressive Forecasting

Brooks Kinch, Xiaozhe Hu, Yilong Huang +6

For autoregressive modeling of chaotic dynamical systems over long time horizons, the stability of both training and inference is a major challenge in building scientific foundatio…

cs.LG2025

Domain-Decomposed Graph Neural Network Surrogate Modeling for Ice Sheets

Adrienne M. Propp, Mauro Perego, Eric C. Cyr +5

Accurate yet efficient surrogate models are essential for large-scale simulations of partial differential equations (PDEs), particularly for uncertainty quantification (UQ) tasks t…

cs.LG2025

Deriving Transformer Architectures as Implicit Multinomial Regression

Jonas A. Actor, Anthony Gruber, Eric C. Cyr

While attention has been empirically shown to improve model performance, it lacks a rigorous mathematical justification. This short paper establishes a novel connection between att…

cs.LG2025

Leveraging KANs for Expedient Training of Multichannel MLPs via Preconditioning and Geometric Refinement

Jonas A. Actor, Graham Harper, Ben Southworth +1

Multilayer perceptrons (MLPs) are a workhorse machine learning architecture, used in a variety of modern deep learning frameworks. However, recently Kolmogorov-Arnold Networks (KAN…