5 papers
OARelatedWork: A Large-Scale Dataset of Related Work Sections with Full-texts from Open Access Sources
Martin Docekal, Martin Fajcik, Pavel Smrz
This paper introduces OARelatedWork: a dataset for related work generation from open-access sources. It is the first large-scale multi-document summarization dataset for related wo…
CzechTopic: A Benchmark for Zero-Shot Topic Localization in Historical Czech Documents
Martin KostelnÃk, Michal HradiÅ¡, Martin DoÄekal
Topic localization aims to identify spans of text that express a given topic defined by a name and description. To study this task, we introduce a human-annotated benchmark based o…
Interoperable verification and dissemination of software assets in repositories using COAR Notify
Matteo Cancellieri, Martin Docekal, David Pride +3
The discoverability, attribution, and reusability of open research software are often hindered by its obscurity within academic manuscripts. To address this, the SoFAIR project (20…
BenCzechMark : A Czech-centric Multitask and Multimetric Benchmark for Large Language Models with Duel Scoring Mechanism
Martin Fajcik, Martin Docekal, Jan Dolezal +15
We present BenCzechMark (BCM), the first comprehensive Czech language benchmark designed for large language models, offering diverse tasks, multiple task formats, and multiple eval…
BiblioPage: A Dataset of Scanned Title Pages for Bibliographic Metadata Extraction
Jan Kohút, Martin DoÄekal, Michal HradiÅ¡ +1
Manual digitization of bibliographic metadata is time consuming and labor intensive, especially for historical and real-world archives with highly variable formatting across docume…