4 papers
Self-Supervised MultiModal Versatile Networks
Jean-Baptiste Alayrac, Adrià Recasens, Rosalia Schneider +6
Videos are a rich source of multi-modal supervision. In this work, we learn representations using self-supervision by leveraging three modalities naturally present in videos: visua…
Data-Efficient Image Recognition with Contrastive Predictive Coding
Olivier J. Hénaff, Aravind Srinivas, Jeffrey De Fauw +4
Human observers can learn to recognize new categories of images from a handful of examples, yet doing so with artificial ones remains an open challenge. We hypothesize that data-ef…
Hierarchical Autoregressive Image Models with Auxiliary Decoders
Jeffrey De Fauw, Sander Dieleman, Karen Simonyan
Autoregressive generative models of images tend to be biased towards capturing local structure, and as a result they often produce samples which are lacking in terms of large-scale…
A Probabilistic U-Net for Segmentation of Ambiguous Images
Simon A. A. Kohl, Bernardino Romera-Paredes, Clemens Meyer +6
Many real-world vision problems suffer from inherent ambiguities. In clinical applications for example, it might not be clear from a CT scan alone which particular region is cancer…