2.9k citations · 3k across the 5 of their papers we have counts for
15 papers
Where Should I Spend My FLOPS? Efficiency Evaluations of Visual Pre-training Methods
Skanda Koppula, Yazhe Li, Evan Shelhamer +5
Self-supervised methods have achieved remarkable success in transfer learning, often achieving the same or better accuracy than supervised pre-training. Most prior work has done so…
Transframer: Arbitrary Frame Prediction with Generative Models
Charlie Nash, João Carreira, Jacob Walker +4
We present a general-purpose framework for image modelling and vision tasks based on probabilistic frame prediction. Our approach unifies a broad range of tasks, from image segment…
Efficient Visual Pretraining with Contrastive Detection
Olivier J. Hénaff, Skanda Koppula, Jean-Baptiste Alayrac +3
Self-supervised pretraining has been shown to yield powerful representations for transfer learning. These performance gains come at a large computational cost however, with state-o…
Perceiver: General Perception with Iterative Attention
Andrew Jaegle, Felix Gimeno, Andrew Brock +3
Biological systems perceive the world by simultaneously processing high-dimensional inputs from modalities as diverse as vision, audition, touch, proprioception, etc. The perceptio…
A Short Note on the Kinetics-700-2020 Human Action Dataset
Lucas Smaira, João Carreira, Eric Noland +3
We describe the 2020 edition of the DeepMind Kinetics human action dataset, which replenishes and extends the Kinetics-700 dataset. In this new version, there are at least 700 vide…
The AVA-Kinetics Localized Human Actions Video Dataset
Ang Li, Meghana Thotakuri, David A. Ross +3
This paper describes the AVA-Kinetics localized human actions video dataset. The dataset is collected by annotating videos from the Kinetics-700 dataset using the AVA annotation pr…