activity
20202022
most citedReinforcement Learning for Weakly Supervised Temporal Grounding of Natural Language in Untrimmed Videos

10 citations · 33 across the 5 of their papers we have counts for

collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2022★ 5 cited

Masked Vision-Language Transformers for Scene Text Recognition

Jie Wu, Ying Peng, Shengming Zhang +2

Scene text recognition (STR) enables computers to recognize and read the text in various real-world scenes. Recent STR models benefit from taking linguistic information in addition…

cs.CV2022

ScalableViT: Rethinking the Context-oriented Generalization of Vision Transformer

Rui Yang, Hailong Ma, Jie Wu +4

The vanilla self-attention mechanism inherently relies on pre-defined and steadfast computational dimensions. Such inflexibility restricts it from possessing context-oriented gener…

cs.CV2021★ 7 cited

Weakly-Supervised Spatio-Temporal Anomaly Detection in Surveillance Video

Jie Wu, Wei Zhang, Guanbin Li +5

In this paper, we introduce a novel task, referred to as Weakly-Supervised Spatio-Temporal Anomaly Detection (WSSTAD) in surveillance video. Specifically, given an untrimmed video,…

cs.CV2020★ 10 cited

Reinforcement Learning for Weakly Supervised Temporal Grounding of Natural Language in Untrimmed Videos

Jie Wu, Guanbin Li, Xiaoguang Han +1

Temporal grounding of natural language in untrimmed videos is a fundamental yet challenging multimedia task facilitating cross-media visual content retrieval. We focus on the weakl…

cs.CV2020★ 8 cited

Fine-Grained Image Captioning with Global-Local Discriminative Objective

Jie Wu, Tianshui Chen, Hefeng Wu +3

Significant progress has been made in recent years in image captioning, an active topic in the fields of vision and language. However, existing methods tend to yield overly general…

cs.CV2020★ 3 cited

Tree-Structured Policy based Progressive Reinforcement Learning for Temporally Language Grounding in Video

Jie Wu, Guanbin Li, Si Liu +1

Temporally language grounding in untrimmed videos is a newly-raised task in video understanding. Most of the existing methods suffer from inferior efficiency, lacking interpretabil…