3.4k citations
- Google (United States)US81 papers
- Google (United Kingdom)GB59 papers
- Massachusetts Institute of TechnologyUS13 papers
- École Normale Supérieure - PSLFR10 papers
- Institut national de recherche en sciences et technologies du numériqueFR10 papers
- University College LondonGB10 papers
- University of OxfordGB10 papers
- University of TorontoCA10 papers
- McGill UniversityCA9 papers
- Imperial College LondonGB8 papers
- University of CambridgeGB8 papers
- Centre de Recherche en InformatiqueFR7 papers
Showing 2023 · cs.CVShow all
2 papers · 2 filters
cs.CV2023
Is ImageNet worth 1 video? Learning strong image encoders from 1 long unlabelled video
Shashanka Venkataramanan, Mamshad Nayeem Rizve, João Carreira +2
Self-supervised learning has unlocked the potential of scaling up pretraining to billions of images, since annotation is unnecessary. But are we making the best use of data? How mo…
cs.CV2023★ 15 cited
Vid2Seq: Large-Scale Pretraining of a Visual Language Model for Dense Video Captioning
Antoine Yang, Arsha Nagrani, Paul Hongsuck Seo +5
In this work, we introduce Vid2Seq, a multi-modal single-stage dense event captioning model pretrained on narrated videos which are readily-available at scale. The Vid2Seq architec…