26 citations · 53 across the 6 of their papers we have counts for
23 papers
Fat-Tailed Variational Inference with Anisotropic Tail Adaptive Flows
Feynman Liang, Liam Hodgkinson, Michael W. Mahoney
While fat-tailed densities commonly arise as posterior and marginal distributions in robust models and scale mixtures, they present challenges when Gaussian-based variational infer…
NoisyMix: Boosting Model Robustness to Common Corruptions
N. Benjamin Erichson, Soon Hoe Lim, Winnie Xu +3
For many real-world applications, obtaining stable and robust statistical performance is more important than simply achieving state-of-the-art predictive test accuracy, and thus ro…
What's Hidden in a One-layer Randomly Weighted Transformer?
Sheng Shen, Zhewei Yao, Douwe Kiela +2
We demonstrate that, hidden within one-layer randomly weighted neural networks, there exist subnetworks that can achieve impressive performance, without ever modifying the weight i…
Stateful ODE-Nets using Basis Function Expansions
Alejandro Queiruga, N. Benjamin Erichson, Liam Hodgkinson +1
The recently-introduced class of ordinary differential equation networks (ODE-Nets) establishes a fruitful connection between deep learning and dynamical systems. In this work, we…
LocalNewton: Reducing Communication Bottleneck for Distributed Learning
Vipul Gupta, Avishek Ghosh, Michal Derezinski +3
To address the communication bottleneck problem in distributed optimization within a master-worker framework, we propose LocalNewton, a distributed second-order algorithm with loca…
ActNN: Reducing Training Memory Footprint via 2-Bit Activation Compressed Training
Jianfei Chen, Lianmin Zheng, Zhewei Yao +4
The increasing size of neural network models has been critical for improvements in their accuracy, but device memory is not growing at the same rate. This creates fundamental chall…