11 citations · 15 across the 4 of their papers we have counts for
4 papers
A Large Cross-Modal Video Retrieval Dataset with Reading Comprehension
Weijia Wu, Yuzhong Zhao, Zhuang Li +4
Most existing cross-modal language-to-video retrieval (VR) research focuses on single-modal input from video, i.e., visual representation, while the text is omnipresent in human en…
FlowText: Synthesizing Realistic Scene Text Video with Optical Flow Estimation
Yuzhong Zhao, Weijia Wu, Zhuang Li +2
Current video text spotting methods can achieve preferable performance, powered with sufficient labeled training data. However, labeling data manually is time-consuming and labor-i…
Real-time End-to-End Video Text Spotter with Contrastive Representation Learning
Wejia Wu, Zhuang Li, Jiahong Li +5
Video text spotting(VTS) is the task that requires simultaneously detecting, tracking and recognizing text in the video. Existing video text spotting methods typically develop soph…
A Bilingual, OpenWorld Video Text Dataset and End-to-end Video Text Spotter with Transformer
Weijia Wu, Yuanqiang Cai, Debing Zhang +5
Most existing video text spotting benchmarks focus on evaluating a single language and scenario with limited data. In this work, we introduce a large-scale, Bilingual, Open World V…