52 citations · 164 across the 13 of their papers we have counts for
16 papers · 1 filter
SDTP: Semantic-aware Decoupled Transformer Pyramid for Dense Image Prediction
Zekun Li, Yufan Liu, Bing Li +3
Although transformer has achieved great progress on computer vision tasks, the scale variation in dense image prediction is still the key challenge. Few effective multi-scale techn…
Channel-wise Topology Refinement Graph Convolution for Skeleton-Based Action Recognition
Yuxin Chen, Ziqi Zhang, Chunfeng Yuan +3
Graph convolutional networks (GCNs) have been widely used and achieved remarkable results in skeleton-based action recognition. In GCNs, graph topology dominates feature aggregatio…
Learn to Match: Automatic Matching Network Design for Visual Tracking
Zhipeng Zhang, Yihao Liu, Xiao Wang +2
Siamese tracking has achieved groundbreaking performance in recent years, where the essence is the efficient matching operator cross-correlation and its variants. Besides the remar…
A Simple and Strong Baseline for Universal Targeted Attacks on Siamese Visual Tracking
Zhenbang Li, Yaya Shi, Jin Gao +4
Siamese trackers are shown to be vulnerable to adversarial attacks recently. However, the existing attack methods craft the perturbations for each video independently, which comes…
Learning to Predict Salient Faces: A Novel Visual-Audio Saliency Model
Yufan Liu, Minglang Qiao, Mai Xu +3
Recently, video streams have occupied a large proportion of Internet traffic, most of which contain human faces. Hence, it is necessary to predict saliency on multiple-face videos,…
Open-book Video Captioning with Retrieve-Copy-Generate Network
Ziqi Zhang, Zhongang Qi, Chunfeng Yuan +4
Due to the rapid emergence of short videos and the requirement for content understanding and creation, the video captioning task has received increasing attention in recent years.…