activity
20172024
most citedSee, Hear, and Read: Deep Aligned Representations

68 citations · 139 across the 14 of their papers we have counts for

collaborators
Showing cs.CVShow all

11 papers · 1 filter

cs.CV2024

A Short Note on Evaluating RepNet for Temporal Repetition Counting in Videos

Debidatta Dwibedi, Yusuf Aytar, Jonathan Tompson +2

We discuss some consistent issues on how RepNet has been evaluated in various papers. As a way to mitigate these issues, we report RepNet performance results on different datasets,…

cs.CV2024

OVR: A Dataset for Open Vocabulary Temporal Repetition Counting in Videos

Debidatta Dwibedi, Yusuf Aytar, Jonathan Tompson +1

We introduce a dataset of annotations of temporal repetitions in videos. The dataset, OVR (pronounced as over), contains annotations for over 72K videos, with each annotation speci…

cs.CV2024

FlexCap: Describe Anything in Images in Controllable Detail

Debidatta Dwibedi, Vidhi Jain, Jonathan Tompson +2

We introduce FlexCap, a vision-language model that generates region-specific descriptions of varying lengths. FlexCap is trained to produce length-conditioned captions for input bo…

cs.CV2023

Learning from One Continuous Video Stream

João Carreira, Michael King, Viorica Pătrăucean +9

We introduce a framework for online learning from a single continuous video stream -- the way people and animals learn, without mini-batches, data augmentation or shuffling. This p…

cs.CV2021

With a Little Help from My Friends: Nearest-Neighbor Contrastive Learning of Visual Representations

Debidatta Dwibedi, Yusuf Aytar, Jonathan Tompson +2

Self-supervised learning algorithms based on instance discrimination train encoders to be invariant to pre-defined transformations of the same instance. While most methods treat di…

cs.CV202012 cited

Large-scale multilingual audio visual dubbing

Yi Yang, Brendan Shillingford, Yannis Assael +9

We describe a system for large-scale audiovisual translation and dubbing, which translates videos from one language to another. The source language's speech content is transcribed…