activity
20202022
most citedReinforcement Learning for Weakly Supervised Temporal Grounding of Natural Language in Untrimmed Videos

10 citations · 33 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CV20225 cited

Masked Vision-Language Transformers for Scene Text Recognition

Jie Wu, Ying Peng, Shengming Zhang +2

Scene text recognition (STR) enables computers to recognize and read the text in various real-world scenes. Recent STR models benefit from taking linguistic information in addition…

cs.CV20217 cited

Weakly-Supervised Spatio-Temporal Anomaly Detection in Surveillance Video

Jie Wu, Wei Zhang, Guanbin Li +5

In this paper, we introduce a novel task, referred to as Weakly-Supervised Spatio-Temporal Anomaly Detection (WSSTAD) in surveillance video. Specifically, given an untrimmed video,…

cs.CV202010 cited

Reinforcement Learning for Weakly Supervised Temporal Grounding of Natural Language in Untrimmed Videos

Jie Wu, Guanbin Li, Xiaoguang Han +1

Temporal grounding of natural language in untrimmed videos is a fundamental yet challenging multimedia task facilitating cross-media visual content retrieval. We focus on the weakl…

cs.CV20208 cited

Fine-Grained Image Captioning with Global-Local Discriminative Objective

Jie Wu, Tianshui Chen, Hefeng Wu +3

Significant progress has been made in recent years in image captioning, an active topic in the fields of vision and language. However, existing methods tend to yield overly general…

cs.CV20203 cited

Tree-Structured Policy based Progressive Reinforcement Learning for Temporally Language Grounding in Video

Jie Wu, Guanbin Li, Si Liu +1

Temporally language grounding in untrimmed videos is a newly-raised task in video understanding. Most of the existing methods suffer from inferior efficiency, lacking interpretabil…