649 citations
- Carnegie Mellon UniversityUS23 papers
- Stanford UniversityUS20 papers
- Google (United States)US14 papers
- Georgia Institute of TechnologyUS13 papers
- Tel Aviv UniversityIL12 papers
- Cornell UniversityUS11 papers
- University of California, BerkeleyUS11 papers
- University College LondonGB10 papers
- Harvard University PressUS9 papers
- Johns Hopkins UniversityUS9 papers
- Massachusetts Institute of TechnologyUS9 papers
- The University of Texas at AustinUS9 papers
4 papers · 2 filters
On the Heavy-Tailed Theory of Stochastic Gradient Descent for Deep Neural Networks
Umut Şimşekli, Mert Gürbüzbalaban, Thanh Huy Nguyen +2
The gradient noise (GN) in the stochastic gradient descent (SGD) algorithm is often considered to be Gaussian in the large data regime by assuming that the \emph{classical} central…
Adaptive Sampling for Estimating Multiple Probability Distributions
Shubhanshu Shekhar, Tara Javidi, Mohammad Ghavamzadeh
We consider the problem of allocating samples to a finite set of discrete distributions in order to learn them uniformly well in terms of four common distance measures: ,…
Bayesian Optimization for Policy Search via Online-Offline Experimentation
Benjamin Letham, Eytan Bakshy
Online field experiments are the gold-standard way of evaluating changes to real-world interactive machine learning systems. Yet our ability to explore complex, multi-dimensional p…
Tensor Variable Elimination for Plated Factor Graphs
Fritz Obermeyer, Eli Bingham, Martin Jankowiak +4
A wide class of machine learning algorithms can be reduced to variable elimination on factor graphs. While factor graphs provide a unifying notation for these algorithms, they do n…