6 citations · 6 across the 1 of their papers we have counts for
2 papers
cs.CV2022★ 6 cited
Learning Audio-Video Modalities from Image Captions
Arsha Nagrani, Paul Hongsuck Seo, Bryan Seybold +4
A major challenge in text-video and text-audio retrieval is the lack of large-scale training data. This is unlike image-captioning, where datasets are in the order of millions of s…
cs.CV2017
PathTrack: Fast Trajectory Annotation with Path Supervision
Santiago Manen, Michael Gygli, Dengxin Dai +1
Progress in Multiple Object Tracking (MOT) has been historically limited by the size of the available datasets. We present an efficient framework to annotate trajectories and use i…