papers

Publications (35)

cs.LG2020

Denoised Smoothing: A Provable Defense for Pretrained Classifiers

Hadi Salman, Mingjie Sun, Greg Yang +2

We present a method for provably defending any pretrained image classifier against adversarial attacks. This method, for instance, allows public vision API providers and u…

cs.NE2017

Lie-Access Neural Turing Machines

Greg Yang, Alexander M. Rush

External neural memory structures have recently become a popular tool for algorithmic deep learning (Graves et al. 2014, Weston et al. 2014). These models generally utilize differe…

cs.LG2021

Tensor Programs IIb: Architectural Universality of Neural Tangent Kernel Training Dynamics

Greg Yang, Etai Littwin

Yang (2020a) recently showed that the Neural Tangent Kernel (NTK) at initialization has an infinite-width limit for a large class of architectures including modern staples such as…

cs.LG2019

Dynamical Isometry and a Mean Field Theory of LSTMs and GRUs

Dar Gilboa, Bo Chang, Minmin Chen +4

Training recurrent neural networks (RNNs) on long sequence tasks is plagued with difficulties arising from the exponential explosion or vanishing of signals as they propagate forwa…

cs.LG2022

Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Greg Yang, Edward J. Hu, Igor Babuschkin +7

Hyperparameter (HP) tuning in deep learning is an expensive process, prohibitively so for neural networks (NNs) with billions of parameters. We show that, in the recently discovere…

cs.NE2021

Tensor Programs III: Neural Matrix Laws

Greg Yang

In a neural network (NN), *weight matrices* linearly transform inputs into *preactivations* that are then transformed nonlinearly into *activations*. A typical NN interleaves multi…