11 citations · 25 across the 9 of their papers we have counts for
9 papers
Visual Information Extraction in the Wild: Practical Dataset and End-to-end Solution
Jianfeng Kuang, Wei Hua, Dingkang Liang +4
Visual information extraction (VIE), which aims to simultaneously perform OCR and information extraction in a unified framework, has drawn increasing attention due to its essential…
The Runner-up Solution for YouTube-VIS Long Video Challenge 2022
Junfeng Wu, Yi Jiang, Qihao Liu +2
This technical report describes our 2nd-place solution for the ECCV 2022 YouTube-VIS Long Video Challenge. We adopt the previously proposed online video instance segmentation metho…
Vision-Language Pre-Training for Boosting Scene Text Detectors
Sibo Song, Jianqiang Wan, Zhibo Yang +4
Recently, vision-language joint representation learning has proven to be highly effective in various scenarios. In this paper, we specifically adapt vision-language joint learning…
An Empirical Study of End-to-End Temporal Action Detection
Xiaolong Liu, Song Bai, Xiang Bai
Temporal action detection (TAD) is an important yet challenging task in video understanding. It aims to simultaneously predict the semantic label and the temporal interval of every…
Few Could Be Better Than All: Feature Sampling and Grouping for Scene Text Detection
Jingqun Tang, Wenqing Zhang, Hongye Liu +4
Recently, transformer-based methods have achieved promising progresses in object detection, as they can eliminate the post-processes like NMS and enrich the deep representations. H…
Comprehensive Benchmark Datasets for Amharic Scene Text Detection and Recognition
Wondimu Dikubab, Dingkang Liang, Minghui Liao +1
Ethiopic/Amharic script is one of the oldest African writing systems, which serves at least 23 languages (e.g., Amharic, Tigrinya) in East Africa for more than 120 million people.…