24 citations · 35 across the 6 of their papers we have counts for
13 papers
Transfer of Representations to Video Label Propagation: Implementation Factors Matter
Daniel McKee, Zitong Zhan, Bing Shuai +3
This work studies feature representations for dense label propagation in video, with a focus on recently proposed methods that learn video correspondence using self-supervised sign…
Multi-Object Tracking with Hallucinated and Unlabeled Videos
Daniel McKee, Bing Shuai, Andrew Berneshawi +4
In this paper, we explore learning end-to-end deep neural trackers without tracking annotations. This is important as large-scale training data is essential for training deep neura…
SiamMOT: Siamese Multi-Object Tracking
Bing Shuai, Andrew Berneshawi, Xinyu Li +2
In this paper, we focus on improving online multi-object tracking (MOT). In particular, we introduce a region-based Siamese Multi-Object Tracking network, which we name SiamMOT. Si…
VidTr: Video Transformer Without Convolutions
Yanyi Zhang, Xinyu Li, Chunhui Liu +6
We introduce Video Transformer (VidTr) with separable-attention for video classification. Comparing with commonly used 3D networks, VidTr is able to aggregate spatio-temporal infor…
NUTA: Non-uniform Temporal Aggregation for Action Recognition
Xinyu Li, Chunhui Liu, Bing Shuai +3
In the world of action recognition research, one primary focus has been on how to construct and train networks to model the spatial-temporal volume of an input video. These methods…
Directional Temporal Modeling for Action Recognition
Xinyu Li, Bing Shuai, Joseph Tighe
Many current activity recognition models use 3D convolutional neural networks (e.g. I3D, I3D-NL) to generate local spatial-temporal features. However, such features do not encode c…