135 citations · 244 across the 9 of their papers we have counts for
9 papers · 1 filter
Learning Multi-Granular Spatio-Temporal Graph Network for Skeleton-based Action Recognition
Tailin Chen, Desen Zhou, Jian Wang +4
The task of skeleton-based action recognition remains a core challenge in human-centred scene understanding due to the multiple granularities and large variation in human motion. E…
Discriminative Latent Semantic Graph for Video Captioning
Yang Bai, Junyan Wang, Yang Long +4
Video captioning aims to automatically generate natural language sentences that can describe the visual contents of a given video. Existing generative models like encoder-decoder f…
Invariant Deep Compressible Covariance Pooling for Aerial Scene Categorization
Shidong Wang, Yi Ren, Gerard Parr +2
Learning discriminative and invariant feature representation is the key to visual image categorization. In this article, we propose a novel invariant deep compressible covariance p…
Query Twice: Dual Mixture Attention Meta Learning for Video Summarization
Junyan Wang, Yang Bai, Yang Long +4
Video summarization aims to select representative frames to retain high-level information, which is usually solved by predicting the segment-wise importance score via a softmax fun…
SOFA-Net: Second-Order and First-order Attention Network for Crowd Counting
Haoran Duan, Shidong Wang, Yu Guan
Automated crowd counting from images/videos has attracted more attention in recent years because of its wide application in smart cities. But modelling the dense crowd heads is cha…
Multi-Granularity Canonical Appearance Pooling for Remote Sensing Scene Classification
S. Wang, Y. Guan, L. Shao
Recognising remote sensing scene images remains challenging due to large visual-semantic discrepancies. These mainly arise due to the lack of detailed annotations that can be emplo…