9 citations · 26 across the 5 of their papers we have counts for
7 papers
Should We Learn Most Likely Functions or Parameters?
Shikai Qiu, Tim G. J. Rudner, Sanyam Kapoor +1
Standard regularized training procedures correspond to maximizing a posterior distribution over parameters, known as maximum a posteriori (MAP) estimation. However, model parameter…
PAC-Bayes Compression Bounds So Tight That They Can Explain Generalization
Sanae Lotfi, Marc Finzi, Sanyam Kapoor +3
While there has been progress in developing non-vacuous generalization bounds for deep neural networks, these bounds tend to be uninformative about why deep learning works. In this…
Pre-Train Your Loss: Easy Bayesian Transfer Learning with Informative Priors
Ravid Shwartz-Ziv, Micah Goldblum, Hossein Souri +4
Deep learning is increasingly moving towards a transfer learning paradigm whereby large foundation models are fine-tuned on downstream tasks, starting from an initialization learne…
On Uncertainty, Tempering, and Data Augmentation in Bayesian Classification
Sanyam Kapoor, Wesley J. Maddox, Pavel Izmailov +1
Aleatoric uncertainty captures the inherent randomness of the data, such as measurement noise. In Bayesian regression, we often use a Gaussian observation model, where we control t…
SKIing on Simplices: Kernel Interpolation on the Permutohedral Lattice for Scalable Gaussian Processes
Sanyam Kapoor, Marc Finzi, Ke Alexander Wang +1
State-of-the-art methods for scalable Gaussian processes use iterative algorithms, requiring fast matrix vector multiplies (MVMs) with the covariance kernel. The Structured Kernel…
First-Order Preconditioning via Hypergradient Descent
Ted Moskovitz, Rui Wang, Janice Lan +4
Standard gradient descent methods are susceptible to a range of issues that can impede training, such as high correlations and different scaling in parameter space.These difficulti…