Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
OARelatedWork: A Large-Scale Dataset of Related Work Sections with Full-texts from Open Access Sources
Martin Docekal, Martin Fajcik, Pavel Smrz
This paper introduces OARelatedWork: a dataset for related work generation from open-access sources. It is the first large-scale multi-document summarization dataset for related wo…
cs.CL2026
CzechTopic: A Benchmark for Zero-Shot Topic Localization in Historical Czech Documents
Martin KostelnÃk, Michal HradiÅ¡, Martin DoÄekal
Topic localization aims to identify spans of text that express a given topic defined by a name and description. To study this task, we introduce a human-annotated benchmark based o…
cs.CL2025
BenCzechMark : A Czech-centric Multitask and Multimetric Benchmark for Large Language Models with Duel Scoring Mechanism
Martin Fajcik, Martin Docekal, Jan Dolezal +15
We present BenCzechMark (BCM), the first comprehensive Czech language benchmark designed for large language models, offering diverse tasks, multiple task formats, and multiple eval…