activity
20172021
most citedA Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark

160 citations · 739 across the 18 of their papers we have counts for

collaborators
Showing stat.MLShow all

6 papers · 1 filter

stat.ML2020

What Do Neural Networks Learn When Trained With Random Labels?

Hartmut Maennel, Ibrahim Alabdulmohsin, Ilya Tolstikhin +4

We study deep neural networks (DNNs) trained on natural image data with entirely random labels. Despite its popularity in the literature, where it is often used to study memorizati…

stat.ML2020

Predicting Neural Network Accuracy from Weights

Thomas Unterthiner, Daniel Keysers, Sylvain Gelly +2

We show experimentally that the accuracy of a trained neural network can be predicted surprisingly well by looking only at its weights, without evaluating it on input data. We moti…

stat.ML202018 cited

On Last-Layer Algorithms for Classification: Decoupling Representation from Uncertainty Estimation

Nicolas Brosse, Carlos Riquelme, Alice Martin +2

Uncertainty quantification for deep learning is a challenging open problem. Bayesian statistics offer a mathematically grounded framework to reason about uncertainties; however, ap…

stat.ML2018

Assessing Generative Models via Precision and Recall

Mehdi S. M. Sajjadi, Olivier Bachem, Mario Lucic +2

Recent advances in generative modeling have led to an increased interest in the study of statistical divergences as means of model comparison. Commonly used evaluation methods, suc…

stat.ML2018

Gradient Descent Quantizes ReLU Network Features

Hartmut Maennel, Olivier Bousquet, Sylvain Gelly

Deep neural networks are often trained in the over-parametrized regime (i.e. with far more parameters than training examples), and understanding why the training converges to solut…

stat.ML201798 cited

From optimal transport to generative modeling: the VEGAN cookbook

Olivier Bousquet, Sylvain Gelly, Ilya Tolstikhin +2

We study unsupervised generative modeling in terms of the optimal transport (OT) problem between true (but unknown) data distribution and the latent variable model distributi…