activity
20182022
most citedDecentralized Federated Averaging

23 citations · 66 across the 12 of their papers we have counts for

collaborators
Showing cs.LGShow all

16 papers · 1 filter

cs.LG2022

glassoformer: a query-sparse transformer for post-fault power grid voltage prediction

Yunling Zheng, Carson Hu, Guang Lin +3

We propose GLassoformer, a novel and efficient transformer architecture leveraging group Lasso regularization to reduce the number of queries of the standard self-attention mechani…

cs.LG2021

How Does Momentum Benefit Deep Neural Networks Architecture Design? A Few Case Studies

Bao Wang, Hedi Xia, Tan Nguyen +1

We present and review an algorithmic and theoretical framework for improving neural network architecture design via momentum. As case studies, we consider how momentum can improve…

cs.LG20215 cited

Training Deep Neural Networks with Adaptive Momentum Inspired by the Quadratic Optimization

Tao Sun, Huaming Ling, Zuoqiang Shi +2

Heavy ball momentum is crucial in accelerating (stochastic) gradient-based optimization algorithms for machine learning. Existing heavy ball momentum is usually weighted by a unifo…

cs.LG202115 cited

Heavy Ball Neural Ordinary Differential Equations

Hedi Xia, Vai Suliafu, Hangjie Ji +4

We propose heavy ball neural ordinary differential equations (HBNODEs), leveraging the continuous limit of the classical momentum accelerated gradient descent, to improve neural OD…

cs.LG20215 cited

FMMformer: Efficient and Flexible Transformer via Decomposed Near-field and Far-field Attention

Tan M. Nguyen, Vai Suliafu, Stanley J. Osher +2

We propose FMMformers, a class of efficient and flexible transformers inspired by the celebrated fast multipole method (FMM) for accelerating interacting particle simulation. FMM d…

cs.LG20211 cited

Robust Certification for Laplace Learning on Geometric Graphs

Matthew Thorpe, Bao Wang

Graph Laplacian (GL)-based semi-supervised learning is one of the most used approaches for classifying nodes in a graph. Understanding and certifying the adversarial robustness of…