14 citations · 36 across the 15 of their papers we have counts for
9 papers · 1 filter
A Short Note on Evaluating RepNet for Temporal Repetition Counting in Videos
Debidatta Dwibedi, Yusuf Aytar, Jonathan Tompson +2
We discuss some consistent issues on how RepNet has been evaluated in various papers. As a way to mitigate these issues, we report RepNet performance results on different datasets,…
OVR: A Dataset for Open Vocabulary Temporal Repetition Counting in Videos
Debidatta Dwibedi, Yusuf Aytar, Jonathan Tompson +1
We introduce a dataset of annotations of temporal repetitions in videos. The dataset, OVR (pronounced as over), contains annotations for over 72K videos, with each annotation speci…
FlexCap: Describe Anything in Images in Controllable Detail
Debidatta Dwibedi, Vidhi Jain, Jonathan Tompson +2
We introduce FlexCap, a vision-language model that generates region-specific descriptions of varying lengths. FlexCap is trained to produce length-conditioned captions for input bo…
With a Little Help from My Friends: Nearest-Neighbor Contrastive Learning of Visual Representations
Debidatta Dwibedi, Yusuf Aytar, Jonathan Tompson +2
Self-supervised learning algorithms based on instance discrimination train encoders to be invariant to pre-defined transformations of the same instance. While most methods treat di…
Counting Out Time: Class Agnostic Video Repetition Counting in the Wild
Debidatta Dwibedi, Yusuf Aytar, Jonathan Tompson +2
We present an approach for estimating the period with which an action is repeated in a video. The crux of the approach lies in constraining the period prediction module to use temp…
An Analysis of Object Representations in Deep Visual Trackers
Ross Goroshin, Jonathan Tompson, Debidatta Dwibedi
Fully convolutional deep correlation networks are integral components of state-of the-art approaches to single object visual tracking. It is commonly assumed that these networks pe…