collaborators

5 papers

cs.CV2025

Archival Faces: Detection of Faces in Digitized Historical Documents

Marek Vaško, Adam Herout, Michal Hradiš

When digitizing historical archives, it is necessary to search for the faces of celebrities and ordinary people, especially in newspapers, link them to the surrounding text, and ma…

cs.CL2025

BenCzechMark : A Czech-centric Multitask and Multimetric Benchmark for Large Language Models with Duel Scoring Mechanism

Martin Fajcik, Martin Docekal, Jan Dolezal +15

We present BenCzechMark (BCM), the first comprehensive Czech language benchmark designed for large language models, offering diverse tasks, multiple task formats, and multiple eval…

cs.CV2025

BiblioPage: A Dataset of Scanned Title Pages for Bibliographic Metadata Extraction

Jan Kohút, Martin Dočekal, Michal Hradiš +1

Manual digitization of bibliographic metadata is time consuming and labor intensive, especially for historical and real-world archives with highly variable formatting across docume…

cs.CV2025

Practical Fine-Tuning of Autoregressive Models on Limited Handwritten Texts

Jan Kohút, Michal Hradiš

A common use case for OCR applications involves users uploading documents and progressively correcting automatic recognition to obtain the final transcript. This correction phase p…

cs.IR2024

A Comparative Study of Text Retrieval Models on DaReCzech

Jakub Stetina, Martin Fajcik, Michal Stefanik +1

This article presents a comprehensive evaluation of 7 off-the-shelf document retrieval models: Splade, Plaid, Plaid-X, SimCSE, Contriever, OpenAI ADA and Gemma2 chosen to determine…