4 papers
AnnoPage Dataset: Dataset of Non-Textual Elements in Documents with Fine-Grained Categorization
Martin Kišš, Michal HradiÅ¡, Martina DvoÅáková +2
We introduce the AnnoPage Dataset, a novel collection of 7,550 pages from historical documents, primarily in Czech and German, spanning from 1485 to the present, focusing on the la…
Towards Writing Style Adaptation in Handwriting Recognition
Jan Kohút, Michal Hradiš, Martin Kišš
One of the challenges of handwriting recognition is to transcribe a large number of vastly different writing styles. State-of-the-art approaches do not explicitly use information a…
Masked Self-Supervised Pre-Training for Text Recognition Transformers on Large-Scale Datasets
Martin Kišš, Michal Hradiš
Self-supervised learning has emerged as a powerful approach for leveraging large-scale unlabeled data to improve model performance in various domains. In this paper, we explore mas…
Self-supervised Pre-training of Text Recognizers
Martin Kišš, Michal Hradiš
In this paper, we investigate self-supervised pre-training methods for document text recognition. Nowadays, large unlabeled datasets can be collected for many research tasks, inclu…