338 citations · 418 across the 28 of their papers we have counts for
51 papers · 1 filter
Understanding Video Scenes through Text: Insights from Text-based Video Question Answering
Soumya Jahagirdar, Minesh Mathew, Dimosthenis Karatzas +1
Researchers have extensively studied the field of vision and language, discovering that both visual and textual content is crucial for understanding scenes effectively. Particularl…
STEP -- Towards Structured Scene-Text Spotting
Sergi Garcia-Bordils, Dimosthenis Karatzas, Marçal Rusiñol
We introduce the structured scene-text spotting task, which requires a scene-text OCR system to spot text in the wild according to a query regular expression. Contrary to generic s…
Reading Between the Lanes: Text VideoQA on the Road
George Tom, Minesh Mathew, Sergi Garcia +2
Text and signs around roads provide crucial information for drivers, vital for safe navigation and situational awareness. Scene text recognition in motion is a challenging problem,…
ICDAR 2023 Competition on Structured Text Extraction from Visually-Rich Document Images
Wenwen Yu, Chengquan Zhang, Haoyu Cao +24
Structured text extraction is one of the most valuable and challenging application directions in the field of Document AI. However, the scenarios of past benchmarks are limited, an…
ICDAR 2023 Video Text Reading Competition for Dense and Small Text
Weijia Wu, Yuzhong Zhao, Zhuang Li +5
Recently, video text detection, tracking, and recognition in natural scenes are becoming very popular in the computer vision community. However, most existing algorithms and benchm…
ICDAR 2023 Competition on Reading the Seal Title
Wenwen Yu, Mingyu Liu, Mingrui Chen +5
Reading seal title text is a challenging task due to the variable shapes of seals, curved text, background noise, and overlapped text. However, this important element is commonly f…