activity
20172026
most citedAutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents

14 citations · 36 across the 15 of their papers we have counts for

collaborators
Showing cs.CVShow all

9 papers · 1 filter

cs.CV2024

A Short Note on Evaluating RepNet for Temporal Repetition Counting in Videos

Debidatta Dwibedi, Yusuf Aytar, Jonathan Tompson +2

We discuss some consistent issues on how RepNet has been evaluated in various papers. As a way to mitigate these issues, we report RepNet performance results on different datasets,…

cs.CV2024

OVR: A Dataset for Open Vocabulary Temporal Repetition Counting in Videos

Debidatta Dwibedi, Yusuf Aytar, Jonathan Tompson +1

We introduce a dataset of annotations of temporal repetitions in videos. The dataset, OVR (pronounced as over), contains annotations for over 72K videos, with each annotation speci…

cs.CV2024★ 1 cited

FlexCap: Describe Anything in Images in Controllable Detail

Debidatta Dwibedi, Vidhi Jain, Jonathan Tompson +2

We introduce FlexCap, a vision-language model that generates region-specific descriptions of varying lengths. FlexCap is trained to produce length-conditioned captions for input bo…

cs.CV2021

With a Little Help from My Friends: Nearest-Neighbor Contrastive Learning of Visual Representations

Debidatta Dwibedi, Yusuf Aytar, Jonathan Tompson +2

Self-supervised learning algorithms based on instance discrimination train encoders to be invariant to pre-defined transformations of the same instance. While most methods treat di…

cs.CV2020★ 3 cited

Counting Out Time: Class Agnostic Video Repetition Counting in the Wild

Debidatta Dwibedi, Yusuf Aytar, Jonathan Tompson +2

We present an approach for estimating the period with which an action is repeated in a video. The crux of the approach lies in constraining the period prediction module to use temp…

cs.CV2020

An Analysis of Object Representations in Deep Visual Trackers

Ross Goroshin, Jonathan Tompson, Debidatta Dwibedi

Fully convolutional deep correlation networks are integral components of state-of the-art approaches to single object visual tracking. It is commonly assumed that these networks pe…