activity
20112024
most citedWavelet Burst Accumulation for turbulence mitigation

27 citations · 49 across the 9 of their papers we have counts for

collaborators
Showing cs.LGShow all

15 papers · 1 filter

cs.LG2021

How Does Momentum Benefit Deep Neural Networks Architecture Design? A Few Case Studies

Bao Wang, Hedi Xia, Tan Nguyen +1

We present and review an algorithmic and theoretical framework for improving neural network architecture design via momentum. As case studies, we consider how momentum can improve…

cs.LG20215 cited

FMMformer: Efficient and Flexible Transformer via Decomposed Near-field and Far-field Attention

Tan M. Nguyen, Vai Suliafu, Stanley J. Osher +2

We propose FMMformers, a class of efficient and flexible transformers inspired by the celebrated fast multipole method (FMM) for accelerating interacting particle simulation. FMM d…

cs.LG2021

Wasserstein Proximal of GANs

Alex Tong Lin, Wuchen Li, Stanley Osher +1

We introduce a new method for training generative adversarial networks by applying the Wasserstein-2 metric proximal on the generators. The approach is based on Wasserstein informa…

cs.LG2020

Sparsity Meets Robustness: Channel Pruning for the Feynman-Kac Formalism Principled Robust Deep Neural Nets

Thu Dinh, Bao Wang, Andrea L. Bertozzi +1

Deep neural nets (DNNs) compression is crucial for adaptation to mobile devices. Though many successful algorithms exist to compress naturally trained DNNs, developing efficient an…

cs.LG2020

Scheduled Restart Momentum for Accelerated Stochastic Gradient Descent

Bao Wang, Tan M. Nguyen, Andrea L. Bertozzi +2

Stochastic gradient descent (SGD) with constant momentum and its variants such as Adam are the optimization algorithms of choice for training deep neural networks (DNNs). Since DNN…

cs.LG2019

Graph Interpolating Activation Improves Both Natural and Robust Accuracies in Data-Efficient Deep Learning

Bao Wang, Stanley J. Osher

Improving the accuracy and robustness of deep neural nets (DNNs) and adapting them to small training data are primary tasks in deep learning research. In this paper, we replace the…