11 citations · 15 across the 3 of their papers we have counts for
6 papers
FACTUAL: A Benchmark for Faithful and Consistent Textual Scene Graph Parsing
Zhuang Li, Yuyang Chai, Terry Yue Zhuo +5
Textual scene graph parsing has become increasingly important in various vision-language applications, including image caption evaluation and image retrieval. However, existing sce…
A Large Cross-Modal Video Retrieval Dataset with Reading Comprehension
Weijia Wu, Yuzhong Zhao, Zhuang Li +4
Most existing cross-modal language-to-video retrieval (VR) research focuses on single-modal input from video, i.e., visual representation, while the text is omnipresent in human en…
FlowText: Synthesizing Realistic Scene Text Video with Optical Flow Estimation
Yuzhong Zhao, Weijia Wu, Zhuang Li +2
Current video text spotting methods can achieve preferable performance, powered with sufficient labeled training data. However, labeling data manually is time-consuming and labor-i…
Real-time End-to-End Video Text Spotter with Contrastive Representation Learning
Wejia Wu, Zhuang Li, Jiahong Li +5
Video text spotting(VTS) is the task that requires simultaneously detecting, tracking and recognizing text in the video. Existing video text spotting methods typically develop soph…
The Third Place Solution for CVPR2022 AVA Accessibility Vision and Autonomy Challenge
Bo Yan, Leilei Cao, Zhuang Li +1
The goal of AVA challenge is to provide vision-based benchmarks and methods relevant to accessibility. In this paper, we introduce the technical details of our submission to the CV…
A Bilingual, OpenWorld Video Text Dataset and End-to-end Video Text Spotter with Transformer
Weijia Wu, Yuanqiang Cai, Debing Zhang +5
Most existing video text spotting benchmarks focus on evaluating a single language and scenario with limited data. In this work, we introduce a large-scale, Bilingual, Open World V…