4 citations · 11 across the 6 of their papers we have counts for
10 papers · 1 filter
Adjacent Leader Decentralized Stochastic Gradient Descent
Haoze He, Jing Wang, Anna Choromanska
This work focuses on the decentralized deep learning optimization framework. We propose Adjacent Leader Decentralized Gradient Descent (AL-DSGD), for improving final model performa…
GRAWA: Gradient-based Weighted Averaging for Distributed Training of Deep Learning Models
Tolga Dimlioglu, Anna Choromanska
We study distributed training of deep learning models in time-constrained environments. We propose a new algorithm that periodically pulls workers towards the center variable compu…
Low-Pass Filtering SGD for Recovering Flat Optima in the Deep Learning Optimization Landscape
Devansh Bisla, Jing Wang, Anna Choromanska
In this paper, we study the sharpness of a deep learning (DL) loss landscape around local minima in order to reveal systematic mechanisms underlying the generalization abilities of…
A Theoretical-Empirical Approach to Estimating Sample Complexity of DNNs
Devansh Bisla, Apoorva Nandini Saridena, Anna Choromanska
This paper focuses on understanding how the generalization error scales with the amount of the training data for deep neural networks (DNNs). Existing techniques in statistical lea…
SGB: Stochastic Gradient Bound Method for Optimizing Partition Functions
Jing Wang, Anna Choromanska
This paper addresses the problem of optimizing partition functions in a stochastic learning setting. We propose a stochastic variant of the bound majorization algorithm that relies…
Learning to Score Behaviors for Guided Policy Optimization
Aldo Pacchiano, Jack Parker-Holder, Yunhao Tang +3
We introduce a new approach for comparing reinforcement learning policies, using Wasserstein distances (WDs) in a newly defined latent behavioral space. We show that by utilizing t…