5 citations · 6 across the 4 of their papers we have counts for
7 papers
Video Referring Expression Comprehension via Transformer with Content-aware Query
Ji Jiang, Meng Cao, Tengtao Song +1
Video Referring Expression Comprehension (REC) aims to localize a target object in video frames referred by the natural language expression. Recently, the Transformerbased methods…
Unsupervised Pre-training for Temporal Action Localization Tasks
Can Zhang, Tianyu Yang, Junwu Weng +3
Unsupervised video representation learning has made remarkable achievements in recent years. However, most existing methods are designed and optimized for video classification. The…
All You Need is a Second Look: Towards Arbitrary-Shaped Text Detection
Meng Cao, Can Zhang, Dongming Yang +1
Arbitrary-shaped text detection is a challenging task since curved texts in the wild are of the complex geometric layouts. Existing mainstream methods follow the instance segmentat…
Video Frame Interpolation via Structure-Motion based Iterative Fusion
Xi Li, Meng Cao, Yingying Tang +4
Video Frame Interpolation synthesizes non-existent images between adjacent frames, with the aim of providing a smooth and consistent visual experience. Two approaches for solving t…
RR-Net: Injecting Interactive Semantics in Human-Object Interaction Detection
Dongming Yang, Yuexian Zou, Can Zhang +2
Human-Object Interaction (HOI) detection devotes to learn how humans interact with surrounding objects. Latest end-to-end HOI detectors are short of relation reasoning, which leads…
CoLA: Weakly-Supervised Temporal Action Localization with Snippet Contrastive Learning
Can Zhang, Meng Cao, Dongming Yang +2
Weakly-supervised temporal action localization (WS-TAL) aims to localize actions in untrimmed videos with only video-level labels. Most existing models follow the "localization by…