activity
20212023
most citedLocVTP: Video-Text Pre-training for Temporal Localization

5 citations · 9 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CV2023

Improving Scene Graph Generation with Superpixel-Based Interaction Learning

Jingyi Wang, Can Zhang, Jinfa Huang +2

Recent advances in Scene Graph Generation (SGG) typically model the relationships among entities utilizing box-level features from pre-defined detectors. We argue that an overlooke…

cs.CV20231 cited

Cross-Modality Time-Variant Relation Learning for Generating Dynamic Scene Graphs

Jingyi Wang, Jinfa Huang, Can Zhang +1

Dynamic scene graphs generated from video clips could help enhance the semantic visual understanding in a wide range of challenging tasks such as environmental perception, autonomo…

cs.CV20225 cited

LocVTP: Video-Text Pre-training for Temporal Localization

Meng Cao, Tianyu Yang, Junwu Weng +3

Video-Text Pre-training (VTP) aims to learn transferable representations for various downstream tasks from large-scale web videos. To date, almost all existing VTP methods are limi…

cs.CV20223 cited

SpatioTemporal Focus for Skeleton-based Action Recognition

Liyu Wu, Can Zhang, Yuexian Zou

Graph convolutional networks (GCNs) are widely adopted in skeleton-based action recognition due to their powerful ability to model data topology. We argue that the performance of r…

cs.CV2021

Long-Short Temporal Modeling for Efficient Action Recognition

Liyu Wu, Yuexian Zou, Can Zhang

Efficient long-short temporal modeling is key for enhancing the performance of action recognition task. In this paper, we propose a new two-stream action recognition network, terme…