activity
20242026
collaborators

5 papers

cs.CL2026

CzechTopic: A Benchmark for Zero-Shot Topic Localization in Historical Czech Documents

Martin Kostelník, Michal Hradiš, Martin Dočekal

Topic localization aims to identify spans of text that express a given topic defined by a name and description. To study this task, we introduce a human-annotated benchmark based o…

cs.CV2025

AnnoPage Dataset: Dataset of Non-Textual Elements in Documents with Fine-Grained Categorization

Martin Kišš, Michal Hradiš, Martina Dvořáková +2

We introduce the AnnoPage Dataset, a novel collection of 7,550 pages from historical documents, primarily in Czech and German, spanning from 1485 to the present, focusing on the la…

cs.CV2025

Masked Self-Supervised Pre-Training for Text Recognition Transformers on Large-Scale Datasets

Martin Kišš, Michal Hradiš

Self-supervised learning has emerged as a powerful approach for leveraging large-scale unlabeled data to improve model performance in various domains. In this paper, we explore mas…

cs.CV2025

TextBite: A Historical Czech Document Dataset for Logical Page Segmentation

Martin Kostelník, Karel Beneš, Michal Hradiš

Logical page segmentation is an important step in document analysis, enabling better semantic representations, information retrieval, and text understanding. Previous approaches de…

cs.CV2024

Self-supervised Pre-training of Text Recognizers

Martin Kišš, Michal Hradiš

In this paper, we investigate self-supervised pre-training methods for document text recognition. Nowadays, large unlabeled datasets can be collected for many research tasks, inclu…