68 citations · 139 across the 14 of their papers we have counts for
11 papers · 1 filter
A Short Note on Evaluating RepNet for Temporal Repetition Counting in Videos
Debidatta Dwibedi, Yusuf Aytar, Jonathan Tompson +2
We discuss some consistent issues on how RepNet has been evaluated in various papers. As a way to mitigate these issues, we report RepNet performance results on different datasets,…
OVR: A Dataset for Open Vocabulary Temporal Repetition Counting in Videos
Debidatta Dwibedi, Yusuf Aytar, Jonathan Tompson +1
We introduce a dataset of annotations of temporal repetitions in videos. The dataset, OVR (pronounced as over), contains annotations for over 72K videos, with each annotation speci…
FlexCap: Describe Anything in Images in Controllable Detail
Debidatta Dwibedi, Vidhi Jain, Jonathan Tompson +2
We introduce FlexCap, a vision-language model that generates region-specific descriptions of varying lengths. FlexCap is trained to produce length-conditioned captions for input bo…
Learning from One Continuous Video Stream
João Carreira, Michael King, Viorica Pătrăucean +9
We introduce a framework for online learning from a single continuous video stream -- the way people and animals learn, without mini-batches, data augmentation or shuffling. This p…
With a Little Help from My Friends: Nearest-Neighbor Contrastive Learning of Visual Representations
Debidatta Dwibedi, Yusuf Aytar, Jonathan Tompson +2
Self-supervised learning algorithms based on instance discrimination train encoders to be invariant to pre-defined transformations of the same instance. While most methods treat di…
Large-scale multilingual audio visual dubbing
Yi Yang, Brendan Shillingford, Yannis Assael +9
We describe a system for large-scale audiovisual translation and dubbing, which translates videos from one language to another. The source language's speech content is transcribed…