5 papers
CzechTopic: A Benchmark for Zero-Shot Topic Localization in Historical Czech Documents
Martin KostelnÃk, Michal HradiÅ¡, Martin DoÄekal
Topic localization aims to identify spans of text that express a given topic defined by a name and description. To study this task, we introduce a human-annotated benchmark based o…
AnnoPage Dataset: Dataset of Non-Textual Elements in Documents with Fine-Grained Categorization
Martin Kišš, Michal HradiÅ¡, Martina DvoÅáková +2
We introduce the AnnoPage Dataset, a novel collection of 7,550 pages from historical documents, primarily in Czech and German, spanning from 1485 to the present, focusing on the la…
Masked Self-Supervised Pre-Training for Text Recognition Transformers on Large-Scale Datasets
Martin Kišš, Michal Hradiš
Self-supervised learning has emerged as a powerful approach for leveraging large-scale unlabeled data to improve model performance in various domains. In this paper, we explore mas…
TextBite: A Historical Czech Document Dataset for Logical Page Segmentation
Martin KostelnÃk, Karel BeneÅ¡, Michal HradiÅ¡
Logical page segmentation is an important step in document analysis, enabling better semantic representations, information retrieval, and text understanding. Previous approaches de…
Self-supervised Pre-training of Text Recognizers
Martin Kišš, Michal Hradiš
In this paper, we investigate self-supervised pre-training methods for document text recognition. Nowadays, large unlabeled datasets can be collected for many research tasks, inclu…