5 papers
Archival Faces: Detection of Faces in Digitized Historical Documents
Marek Vaško, Adam Herout, Michal Hradiš
When digitizing historical archives, it is necessary to search for the faces of celebrities and ordinary people, especially in newspapers, link them to the surrounding text, and ma…
BenCzechMark : A Czech-centric Multitask and Multimetric Benchmark for Large Language Models with Duel Scoring Mechanism
Martin Fajcik, Martin Docekal, Jan Dolezal +15
We present BenCzechMark (BCM), the first comprehensive Czech language benchmark designed for large language models, offering diverse tasks, multiple task formats, and multiple eval…
BiblioPage: A Dataset of Scanned Title Pages for Bibliographic Metadata Extraction
Jan Kohút, Martin DoÄekal, Michal HradiÅ¡ +1
Manual digitization of bibliographic metadata is time consuming and labor intensive, especially for historical and real-world archives with highly variable formatting across docume…
Practical Fine-Tuning of Autoregressive Models on Limited Handwritten Texts
Jan Kohút, Michal Hradiš
A common use case for OCR applications involves users uploading documents and progressively correcting automatic recognition to obtain the final transcript. This correction phase p…
A Comparative Study of Text Retrieval Models on DaReCzech
Jakub Stetina, Martin Fajcik, Michal Stefanik +1
This article presents a comprehensive evaluation of 7 off-the-shelf document retrieval models: Splade, Plaid, Plaid-X, SimCSE, Contriever, OpenAI ADA and Gemma2 chosen to determine…