activity
20132026
most citedScore-Based Generative Modeling through Stochastic Differential Equations

1.3k citations · 2.5k across the 28 of their papers we have counts for

collaborators
Showing stat.MLShow all

12 papers · 1 filter

stat.ML202029 cited

Infinite attention: NNGP and NTK for deep attention networks

Jiri Hron, Yasaman Bahri, Jascha Sohl-Dickstein +1

There is a growing amount of literature on the relationship between wide neural networks (NNs) and Gaussian processes (GPs), identifying an equivalence between the two for a variet…

stat.ML2020

Exact posterior distributions of wide Bayesian neural networks

Jiri Hron, Yasaman Bahri, Roman Novak +2

Recent work has shown that the prior over functions induced by a deep Bayesian neural network (BNN) behaves as a Gaussian process (GP) as the width of all layers becomes large. How…

stat.ML202058 cited

The large learning rate phase of deep learning: the catapult mechanism

Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer +2

The choice of initial learning rate can have a profound effect on the performance of deep networks. We present a class of neural networks with solvable training dynamics, and confi…

stat.ML201957 cited

Neural Tangents: Fast and Easy Infinite Neural Networks in Python

Roman Novak, Lechao Xiao, Jiri Hron +4

Neural Tangents is a library designed to enable research into infinite-width neural networks. It provides a high-level API for specifying complex and hierarchical neural network ar…

stat.ML2019

Wide Neural Networks of Any Depth Evolve as Linear Models Under Gradient Descent

Jaehoon Lee, Lechao Xiao, Samuel S. Schoenholz +4

A longstanding goal in deep learning research has been to precisely characterize training and generalization. However, the often complex loss landscapes of neural networks have mad…

stat.ML20193 cited

Eliminating all bad Local Minima from Loss Landscapes without even adding an Extra Unit

Jascha Sohl-Dickstein, Kenji Kawaguchi

Recent work has noted that all bad local minima can be removed from neural network loss landscapes, by adding a single unit with a particular parameterization. We show that the cor…