activity
20182022
most citedChannel-wise Topology Refinement Graph Convolution for Skeleton-Based Action Recognition

52 citations · 98 across the 5 of their papers we have counts for

collaborators

6 papers

cs.CV20225 cited

CREATE: A Benchmark for Chinese Short Video Retrieval and Title Generation

Ziqi Zhang, Yuxin Chen, Zongyang Ma +5

Previous works of video captioning aim to objectively describe the video's actual content, which lacks subjective and attractive expression, limiting its practical application scen…

cs.CV202152 cited

Channel-wise Topology Refinement Graph Convolution for Skeleton-Based Action Recognition

Yuxin Chen, Ziqi Zhang, Chunfeng Yuan +3

Graph convolutional networks (GCNs) have been widely used and achieved remarkable results in skeleton-based action recognition. In GCNs, graph topology dominates feature aggregatio…

cs.CV202039 cited

Object Relational Graph with Teacher-Recommended Learning for Video Captioning

Ziqi Zhang, Yaya Shi, Chunfeng Yuan +4

Taking full advantage of the information from both vision and language is critical for the video captioning task. Existing models lack adequate visual representation due to the neg…

cs.CL2019

VATEX Captioning Challenge 2019: Multi-modal Information Fusion and Multi-stage Training Strategy for Video Captioning

Ziqi Zhang, Yaya Shi, Jiutong Wei +3

Multi-modal information is essential to describe what has happened in a video. In this work, we represent videos by various appearance, motion and audio information guided with vid…

cs.CV20192 cited

Multimodal Semantic Attention Network for Video Captioning

Liang Sun, Bing Li, Chunfeng Yuan +2

Inspired by the fact that different modalities in videos carry complementary information, we propose a Multimodal Semantic Attention Network(MSAN), which is a new encoder-decoder f…

cs.CV2018

Interaction-aware Spatio-temporal Pyramid Attention Networks for Action Classification

Yang Du, Chunfeng Yuan, Bing Li +3

Local features at neighboring spatial positions in feature maps have high correlation since their receptive fields are often overlapped. Self-attention usually uses the weighted su…