115 citations · 158 across the 20 of their papers we have counts for
7 papers · 1 filter
NUTA: Non-uniform Temporal Aggregation for Action Recognition
Xinyu Li, Chunhui Liu, Bing Shuai +3
In the world of action recognition research, one primary focus has been on how to construct and train networks to model the spatial-temporal volume of an input video. These methods…
A Comprehensive Study of Deep Video Action Recognition
Yi Zhu, Xinyu Li, Chunhui Liu +7
Video action recognition is one of the representative tasks for video understanding. Over the last decade, we have witnessed great advancements in video action recognition thanks t…
Directional Temporal Modeling for Action Recognition
Xinyu Li, Bing Shuai, Joseph Tighe
Many current activity recognition models use 3D convolutional neural networks (e.g. I3D, I3D-NL) to generate local spatial-temporal features. However, such features do not encode c…
Multi-Object Tracking with Siamese Track-RCNN
Bing Shuai, Andrew G. Berneshawi, Davide Modolo +1
Multi-object tracking systems often consist of a combination of a detector, a short term linker, a re-identification feature extractor and a solver that takes the output from these…
Understanding the impact of mistakes on background regions in crowd counting
Davide Modolo, Bing Shuai, Rahul Rama Varior +1
Every crowd counting researcher has likely observed their model output wrong positive predictions on image regions not containing any person. But how often do these mistakes happen…
Combining detection and tracking for human pose estimation in videos
Manchen Wang, Joseph Tighe, Davide Modolo
We propose a novel top-down approach that tackles the problem of multi-person human pose estimation and tracking in videos. In contrast to existing top-down approaches, our method…