most citedA Bilingual, OpenWorld Video Text Dataset and End-to-end Video Text Spotter with Transformer

11 citations · 15 across the 3 of their papers we have counts for

collaborators

6 papers

cs.CL2023

FACTUAL: A Benchmark for Faithful and Consistent Textual Scene Graph Parsing

Zhuang Li, Yuyang Chai, Terry Yue Zhuo +5

Textual scene graph parsing has become increasingly important in various vision-language applications, including image caption evaluation and image retrieval. However, existing sce…

cs.CV2023

A Large Cross-Modal Video Retrieval Dataset with Reading Comprehension

Weijia Wu, Yuzhong Zhao, Zhuang Li +4

Most existing cross-modal language-to-video retrieval (VR) research focuses on single-modal input from video, i.e., visual representation, while the text is omnipresent in human en…

cs.CV2023

FlowText: Synthesizing Realistic Scene Text Video with Optical Flow Estimation

Yuzhong Zhao, Weijia Wu, Zhuang Li +2

Current video text spotting methods can achieve preferable performance, powered with sufficient labeled training data. However, labeling data manually is time-consuming and labor-i…

cs.CV20224 cited

Real-time End-to-End Video Text Spotter with Contrastive Representation Learning

Wejia Wu, Zhuang Li, Jiahong Li +5

Video text spotting(VTS) is the task that requires simultaneously detecting, tracking and recognizing text in the video. Existing video text spotting methods typically develop soph…

cs.CV2022

The Third Place Solution for CVPR2022 AVA Accessibility Vision and Autonomy Challenge

Bo Yan, Leilei Cao, Zhuang Li +1

The goal of AVA challenge is to provide vision-based benchmarks and methods relevant to accessibility. In this paper, we introduce the technical details of our submission to the CV…

cs.CV202111 cited

A Bilingual, OpenWorld Video Text Dataset and End-to-end Video Text Spotter with Transformer

Weijia Wu, Yuanqiang Cai, Debing Zhang +5

Most existing video text spotting benchmarks focus on evaluating a single language and scenario with limited data. In this work, we introduce a large-scale, Bilingual, Open World V…