23 citations · 66 across the 12 of their papers we have counts for
16 papers · 1 filter
glassoformer: a query-sparse transformer for post-fault power grid voltage prediction
Yunling Zheng, Carson Hu, Guang Lin +3
We propose GLassoformer, a novel and efficient transformer architecture leveraging group Lasso regularization to reduce the number of queries of the standard self-attention mechani…
How Does Momentum Benefit Deep Neural Networks Architecture Design? A Few Case Studies
Bao Wang, Hedi Xia, Tan Nguyen +1
We present and review an algorithmic and theoretical framework for improving neural network architecture design via momentum. As case studies, we consider how momentum can improve…
Training Deep Neural Networks with Adaptive Momentum Inspired by the Quadratic Optimization
Tao Sun, Huaming Ling, Zuoqiang Shi +2
Heavy ball momentum is crucial in accelerating (stochastic) gradient-based optimization algorithms for machine learning. Existing heavy ball momentum is usually weighted by a unifo…
Heavy Ball Neural Ordinary Differential Equations
Hedi Xia, Vai Suliafu, Hangjie Ji +4
We propose heavy ball neural ordinary differential equations (HBNODEs), leveraging the continuous limit of the classical momentum accelerated gradient descent, to improve neural OD…
FMMformer: Efficient and Flexible Transformer via Decomposed Near-field and Far-field Attention
Tan M. Nguyen, Vai Suliafu, Stanley J. Osher +2
We propose FMMformers, a class of efficient and flexible transformers inspired by the celebrated fast multipole method (FMM) for accelerating interacting particle simulation. FMM d…
Robust Certification for Laplace Learning on Geometric Graphs
Matthew Thorpe, Bao Wang
Graph Laplacian (GL)-based semi-supervised learning is one of the most used approaches for classifying nodes in a graph. Understanding and certifying the adversarial robustness of…