115 citations · 238 across the 24 of their papers we have counts for
29 papers
Hybrid Relation Guided Set Matching for Few-shot Action Recognition
Xiang Wang, Shiwei Zhang, Zhiwu Qing +5
Current few-shot action recognition methods reach impressive performance by learning discriminative features for each video via episodic training and designing various temporal ali…
Learning from Untrimmed Videos: Self-Supervised Video Representation Learning with Hierarchical Consistency
Zhiwu Qing, Shiwei Zhang, Ziyuan Huang +6
Natural videos provide rich visual contents for self-supervised learning. Yet most existing approaches for learning spatio-temporal representations rely on manually trimmed videos,…
Style Transformer for Image Inversion and Editing
Xueqi Hu, Qiusheng Huang, Zhengyi Shi +4
Existing GAN inversion methods fail to provide latent codes for reliable reconstruction and flexible editing simultaneously. This paper presents a transformer-based image inversion…
CondNet: Conditional Classifier for Scene Segmentation
Changqian Yu, Yuanjie Shao, Changxin Gao +1
The fully convolutional network (FCN) has achieved tremendous success in dense visual recognition tasks, such as scene segmentation. The last layer of FCN is typically a global cla…
Weakly Supervised Person Search with Region Siamese Networks
Chuchu Han, Kai Su, Dongdong Yu +5
Supervised learning is dominant in person search, but it requires elaborate labeling of bounding boxes and identities. Large-scale labeled training data is often difficult to colle…
Exploring Stronger Feature for Temporal Action Localization
Zhiwu Qing, Xiang Wang, Ziyuan Huang +6
Temporal action localization aims to localize starting and ending time with action category. Limited by GPU memory, mainstream methods pre-extract features for each video. Therefor…