6 citations · 15 across the 18 of their papers we have counts for
6 papers · 1 filter
Only Time Can Tell: Discovering Temporal Data for Temporal Modeling
Laura Sevilla-Lara, Shengxin Zha, Zhicheng Yan +3
Understanding temporal information and how the visual world changes over time is a fundamental ability of intelligent systems. In video understanding, temporal information is at th…
FASTER Recurrent Networks for Efficient Video Classification
Linchao Zhu, Laura Sevilla-Lara, Du Tran +3
Typical video classification methods often divide a video into short clips, do inference on each clip independently, then aggregate the clip-level predictions to generate the video…
Video Modeling with Correlation Networks
Heng Wang, Du Tran, Lorenzo Torresani +1
Motion is a salient cue to recognize actions in video. Modern action recognition models leverage motion information either explicitly by using optical flow as input or implicitly b…
Large-scale weakly-supervised pre-training for video action recognition
Deepti Ghadiyaram, Matt Feiszli, Du Tran +3
Current fully-supervised video datasets consist of only a few hundred thousand videos and fewer than a thousand domain-specific labels. This hinders the progress towards advanced v…
What Makes Training Multi-Modal Classification Networks Hard?
Weiyao Wang, Du Tran, Matt Feiszli
Consider end-to-end training of a multi-modal vs. a single-modal network on a task with multiple input modalities: the multi-modal network receives more information, so it should m…
Video Classification with Channel-Separated Convolutional Networks
Du Tran, Heng Wang, Lorenzo Torresani +1
Group convolution has been shown to offer great computational savings in various 2D convolutional architectures for image classification. It is natural to ask: 1) if group convolut…