output
20142025
most citedBootstrap your own latent: A new approach to self-supervised Learning

3.4k citations

Showing cs.CVShow all

12 papers · 1 filter

cs.CV202315 cited

Vid2Seq: Large-Scale Pretraining of a Visual Language Model for Dense Video Captioning

Antoine Yang, Arsha Nagrani, Paul Hongsuck Seo +5

In this work, we introduce Vid2Seq, a multi-modal single-stage dense event captioning model pretrained on narrated videos which are readily-available at scale. The Vid2Seq architec…

cs.CV20228 cited

AlignSDF: Pose-Aligned Signed Distance Fields for Hand-Object Reconstruction

Zerui Chen, Yana Hasson, Cordelia Schmid +1

Recent work achieved impressive progress towards joint reconstruction of hands and manipulated objects from monocular color images. Existing methods focus on two alternative repres…

cs.CV20219 cited

Generating Images with Sparse Representations

Charlie Nash, Jacob Menick, Sander Dieleman +1

The high dimensionality of images presents architecture and sampling-efficiency challenges for likelihood-based generative models. Previous approaches such as VQ-VAE use deep autoe…

cs.CV202125 cited

Predicting Video with VQVAE

Jacob Walker, Ali Razavi, Aäron van den Oord

In recent years, the task of video prediction-forecasting future video given past video frames-has attracted attention in the research community. In this paper we propose a novel a…

cs.CV2021256 cited

High-Performance Large-Scale Image Recognition Without Normalization

Andrew Brock, Soham De, Samuel L. Smith +1

Batch normalization is a key component of most image classification models, but it has many undesirable properties stemming from its dependence on the batch size and interactions b…

cs.CV20201 cited

Adaptive Text Recognition through Visual Matching

Chuhan Zhang, Ankush Gupta, Andrew Zisserman

In this work, our objective is to address the problems of generalization and flexibility for text recognition in documents. We introduce a new model that exploits the repetitive na…