activity
20192026
most citedOptimization Theory for ReLU Neural Networks Trained with Normalization Layers

13 citations · 14 across the 5 of their papers we have counts for

collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG2025

FarSkip-Collective: Unhobbling Blocking Communication in Mixture of Experts Models

Yonatan Dukler, Guihong Li, Deval Shah +3

Blocking communication presents a major hurdle in running MoEs efficiently in distributed settings. To address this, we present FarSkip-Collective which modifies the architecture o…

cs.LG20241 cited

B'MOJO: Hybrid State Space Realizations of Foundation Models with Eidetic and Fading Memory

Luca Zancato, Arjun Seshadri, Yonatan Dukler +6

We describe a family of architectures to support transductive inference by allowing memory to grow to a finite but a-priori unknown bound while making efficient use of finite resou…

cs.LG2023

SAFE: Machine Unlearning With Shard Graphs

Yonatan Dukler, Benjamin Bowman, Alessandro Achille +3

We present Synergy Aware Forgetting Ensemble (SAFE), a method to adapt large models on a diverse collection of data while minimizing the expected cost to remove the influence of tr…

cs.LG2023

Your representations are in the network: composable and parallel adaptation for large scale models

Yonatan Dukler, Alessandro Achille, Hao Yang +7

We propose InCA, a lightweight method for transfer learning that cross-attends to any activation layer of a pre-trained model. During training, InCA uses a single forward pass to e…

cs.LG202013 cited

Optimization Theory for ReLU Neural Networks Trained with Normalization Layers

Yonatan Dukler, Quanquan Gu, Guido Montúfar

The success of deep neural networks is in part due to the use of normalization layers. Normalization layers like Batch Normalization, Layer Normalization and Weight Normalization a…

cs.LG2019

Wasserstein Diffusion Tikhonov Regularization

Alex Tong Lin, Yonatan Dukler, Wuchen Li +1

We propose regularization strategies for learning discriminative models that are robust to in-class variations of the input data. We use the Wasserstein-2 geometry to capture seman…