3.4k citations
- Google (United States)US56 papers
- Google (United Kingdom)GB17 papers
- University of TorontoCA7 papers
- Centre de Recherche en InformatiqueFR6 papers
- Centre de Recherche en Informatique, Signal et Automatique de LilleFR6 papers
- University of AlbertaCA6 papers
- University of OxfordGB6 papers
- Carnegie Mellon UniversityUS5 papers
- Columbia UniversityUS5 papers
- McGill UniversityCA5 papers
- Afterschool AllianceUS4 papers
- École Normale Supérieure - PSLFR4 papers
12 papers · 1 filter
Vid2Seq: Large-Scale Pretraining of a Visual Language Model for Dense Video Captioning
Antoine Yang, Arsha Nagrani, Paul Hongsuck Seo +5
In this work, we introduce Vid2Seq, a multi-modal single-stage dense event captioning model pretrained on narrated videos which are readily-available at scale. The Vid2Seq architec…
AlignSDF: Pose-Aligned Signed Distance Fields for Hand-Object Reconstruction
Zerui Chen, Yana Hasson, Cordelia Schmid +1
Recent work achieved impressive progress towards joint reconstruction of hands and manipulated objects from monocular color images. Existing methods focus on two alternative repres…
Generating Images with Sparse Representations
Charlie Nash, Jacob Menick, Sander Dieleman +1
The high dimensionality of images presents architecture and sampling-efficiency challenges for likelihood-based generative models. Previous approaches such as VQ-VAE use deep autoe…
Predicting Video with VQVAE
Jacob Walker, Ali Razavi, Aäron van den Oord
In recent years, the task of video prediction-forecasting future video given past video frames-has attracted attention in the research community. In this paper we propose a novel a…
High-Performance Large-Scale Image Recognition Without Normalization
Andrew Brock, Soham De, Samuel L. Smith +1
Batch normalization is a key component of most image classification models, but it has many undesirable properties stemming from its dependence on the batch size and interactions b…
Adaptive Text Recognition through Visual Matching
Chuhan Zhang, Ankush Gupta, Andrew Zisserman
In this work, our objective is to address the problems of generalization and flexibility for text recognition in documents. We introduce a new model that exploits the repetitive na…