148 citations · 265 across the 4 of their papers we have counts for
7 papers
SSCAP: Self-supervised Co-occurrence Action Parsing for Unsupervised Temporal Action Segmentation
Zhe Wang, Hao Chen, Xinyu Li +4
Temporal action segmentation is a task to classify each frame in the video with an action label. However, it is quite expensive to annotate every frame in a large corpus of videos…
VidTr: Video Transformer Without Convolutions
Yanyi Zhang, Xinyu Li, Chunhui Liu +6
We introduce Video Transformer (VidTr) with separable-attention for video classification. Comparing with commonly used 3D networks, VidTr is able to aggregate spatio-temporal infor…
Selective Feature Compression for Efficient Activity Recognition Inference
Chunhui Liu, Xinyu Li, Hao Chen +2
Most action recognition solutions rely on dense sampling to precisely cover the informative temporal clip. Extensively searching temporal region is expensive for a real-world appli…
NUTA: Non-uniform Temporal Aggregation for Action Recognition
Xinyu Li, Chunhui Liu, Bing Shuai +3
In the world of action recognition research, one primary focus has been on how to construct and train networks to model the spatial-temporal volume of an input video. These methods…
A Comprehensive Study of Deep Video Action Recognition
Yi Zhu, Xinyu Li, Chunhui Liu +7
Video action recognition is one of the representative tasks for video understanding. Over the last decade, we have witnessed great advancements in video action recognition thanks t…
Triplet Online Instance Matching Loss for Person Re-identification
Ye Li, Guangqiang Yin, Chunhui Liu +2
Mining the shared features of same identity in different scene, and the unique features of different identity in same scene, are most significant challenges in the field of person…