11 citations · 20 across the 9 of their papers we have counts for
20 papers · 1 filter
MorfFlex: Handling Rich Morphology
Jaroslava Hlaváčová, Marie Mikulová, Barbora Štěpánková +2
We present MorfFlex, a morphological dictionary architecture suitable for languages with extensive regularity in both inflection and derivation. As the primary example of MorfFlex…
Meet UD_Czech-PDTC: A Large and Genre-Rich Treebank in Universal Dependencies
Marie Mikulová, Barbora Štěpánková, Daniel Zeman +3
Czech has been part of Universal Dependencies since its first release in 2015. It has also been one of the best represented languages, with the Prague Dependency Treebank being ord…
Prague Dependency Treebank -- Consolidated 2.0: Enriching a Complex Annotation Scheme
Marie Mikulová, Jiří Mírovský, Milan Straka +4
The Prague Dependency Treebank framework is unique in its attempt to systematically include and link different layers of language, including a meaning representation with several t…
Refining Czech GEC: Insights from a Multi-Experiment Approach
Petr Pechman, Milan Straka, Jana Straková +1
We present a grammar error correction (GEC) system that achieves state of the art for the Czech language. Our system is based on a neural network translation approach with the Tran…
Quality and Efficiency of Manual Annotation: Pre-annotation Bias
Marie Mikulová, Milan Straka, Jan Štěpánek +2
This paper presents an analysis of annotation using an automatic pre-annotation for a mid-level annotation complexity task -- dependency syntax annotation. It compares the annotati…
DaMuEL: A Large Multilingual Dataset for Entity Linking
David Kubeša, Milan Straka
We present DaMuEL, a large Multilingual Dataset for Entity Linking containing data in 53 languages. DaMuEL consists of two components: a knowledge base that contains language-agnos…