348 citations · 350 across the 2 of their papers we have counts for
3 papers
cs.CV2019★ 2 cited
UniDual: A Unified Model for Image and Video Understanding
Yufei Wang, Du Tran, Lorenzo Torresani
Although a video is effectively a sequence of images, visual perception systems typically model images and videos separately, thus failing to exploit the correlation and the synerg…
cs.CV2019
Large-scale weakly-supervised pre-training for video action recognition
Deepti Ghadiyaram, Matt Feiszli, Du Tran +3
Current fully-supervised video datasets consist of only a few hundred thousand videos and fewer than a thousand domain-specific labels. This hinders the progress towards advanced v…
cs.CV2017★ 348 cited
ConvNet Architecture Search for Spatiotemporal Feature Learning
Du Tran, Jamie Ray, Zheng Shou +2
Learning image representations with ConvNets by pre-training on ImageNet has proven useful across many visual understanding tasks including object detection, semantic segmentation,…