27 citations · 49 across the 9 of their papers we have counts for
15 papers · 1 filter
How Does Momentum Benefit Deep Neural Networks Architecture Design? A Few Case Studies
Bao Wang, Hedi Xia, Tan Nguyen +1
We present and review an algorithmic and theoretical framework for improving neural network architecture design via momentum. As case studies, we consider how momentum can improve…
FMMformer: Efficient and Flexible Transformer via Decomposed Near-field and Far-field Attention
Tan M. Nguyen, Vai Suliafu, Stanley J. Osher +2
We propose FMMformers, a class of efficient and flexible transformers inspired by the celebrated fast multipole method (FMM) for accelerating interacting particle simulation. FMM d…
Wasserstein Proximal of GANs
Alex Tong Lin, Wuchen Li, Stanley Osher +1
We introduce a new method for training generative adversarial networks by applying the Wasserstein-2 metric proximal on the generators. The approach is based on Wasserstein informa…
Sparsity Meets Robustness: Channel Pruning for the Feynman-Kac Formalism Principled Robust Deep Neural Nets
Thu Dinh, Bao Wang, Andrea L. Bertozzi +1
Deep neural nets (DNNs) compression is crucial for adaptation to mobile devices. Though many successful algorithms exist to compress naturally trained DNNs, developing efficient an…
Scheduled Restart Momentum for Accelerated Stochastic Gradient Descent
Bao Wang, Tan M. Nguyen, Andrea L. Bertozzi +2
Stochastic gradient descent (SGD) with constant momentum and its variants such as Adam are the optimization algorithms of choice for training deep neural networks (DNNs). Since DNN…
Graph Interpolating Activation Improves Both Natural and Robust Accuracies in Data-Efficient Deep Learning
Bao Wang, Stanley J. Osher
Improving the accuracy and robustness of deep neural nets (DNNs) and adapting them to small training data are primary tasks in deep learning research. In this paper, we replace the…