5 citations · 9 across the 5 of their papers we have counts for
5 papers
Improving Scene Graph Generation with Superpixel-Based Interaction Learning
Jingyi Wang, Can Zhang, Jinfa Huang +2
Recent advances in Scene Graph Generation (SGG) typically model the relationships among entities utilizing box-level features from pre-defined detectors. We argue that an overlooke…
Cross-Modality Time-Variant Relation Learning for Generating Dynamic Scene Graphs
Jingyi Wang, Jinfa Huang, Can Zhang +1
Dynamic scene graphs generated from video clips could help enhance the semantic visual understanding in a wide range of challenging tasks such as environmental perception, autonomo…
LocVTP: Video-Text Pre-training for Temporal Localization
Meng Cao, Tianyu Yang, Junwu Weng +3
Video-Text Pre-training (VTP) aims to learn transferable representations for various downstream tasks from large-scale web videos. To date, almost all existing VTP methods are limi…
SpatioTemporal Focus for Skeleton-based Action Recognition
Liyu Wu, Can Zhang, Yuexian Zou
Graph convolutional networks (GCNs) are widely adopted in skeleton-based action recognition due to their powerful ability to model data topology. We argue that the performance of r…
Long-Short Temporal Modeling for Efficient Action Recognition
Liyu Wu, Yuexian Zou, Can Zhang
Efficient long-short temporal modeling is key for enhancing the performance of action recognition task. In this paper, we propose a new two-stream action recognition network, terme…