activity
20122021
most citedA simpler approach to obtaining an O(1/t) convergence rate for the projected stochastic subgradient method

33 citations · 113 across the 8 of their papers we have counts for

collaborators
Showing 2018Show all

6 papers · 1 filter

cs.LG2018

SLANG: Fast Structured Covariance Approximations for Bayesian Deep Learning with Natural Gradient

Aaron Mishkin, Frederik Kunstner, Didrik Nielsen +2

Uncertainty estimation in large deep-learning models is a computationally challenging task, where it is difficult to form even a Gaussian approximation to the posterior distributio…

cs.LG2018

Fast and Faster Convergence of SGD for Over-Parameterized Models and an Accelerated Perceptron

Sharan Vaswani, Francis Bach, Mark Schmidt

Modern machine learning focuses on highly expressive models that are able to fit or interpolate the data completely, resulting in zero training loss. For such models, we show that…

cs.LG2018

Combining Bayesian Optimization and Lipschitz Optimization

Mohamed Osama Ahmed, Sharan Vaswani, Mark Schmidt

Bayesian optimization and Lipschitz optimization have developed alternative techniques for optimizing black-box functions. They each exploit a different form of prior about the fun…

cs.LG2018

A Less Biased Evaluation of Out-of-distribution Sample Detectors

Alireza Shafaei, Mark Schmidt, James J. Little

In the real world, a learning system could receive an input that is unlike anything it has seen during training. Unfortunately, out-of-distribution samples can lead to unpredictabl…

cs.CV2018

Where are the Blobs: Counting by Localization with Point Supervision

Issam H. Laradji, Negar Rostamzadeh, Pedro O. Pinheiro +2

Object counting is an important task in computer vision due to its growing demand in applications such as surveillance, traffic monitoring, and counting everyday objects. State-of-…

cs.LG2018

New Insights into Bootstrapping for Bandits

Sharan Vaswani, Branislav Kveton, Zheng Wen +3

We investigate the use of bootstrapping in the bandit setting. We first show that the commonly used non-parametric bootstrapping (NPB) procedure can be provably inefficient and est…