activity
20182023
most citedPrague Dependency Treebank -- Consolidated 1.0

11 citations · 20 across the 6 of their papers we have counts for

collaborators

16 papers

cs.CL2023

Quality and Efficiency of Manual Annotation: Pre-annotation Bias

Marie Mikulová, Milan Straka, Jan Štěpánek +2

This paper presents an analysis of annotation using an automatic pre-annotation for a mid-level annotation complexity task -- dependency syntax annotation. It compares the annotati…

cs.CL20231 cited

DaMuEL: A Large Multilingual Dataset for Entity Linking

David Kubeša, Milan Straka

We present DaMuEL, a large Multilingual Dataset for Entity Linking containing data in 53 languages. DaMuEL consists of two components: a knowledge base that contains language-agnos…

cs.CL20211 cited

Diacritics Restoration using BERT with Analysis on Czech language

Jakub Náplava, Milan Straka, Jana Straková

We propose a new architecture for diacritics restoration based on contextualized embeddings, namely BERT, and we evaluate it on 12 languages with diacritics. Furthermore, we conduc…

cs.CL2021

RobeCzech: Czech RoBERTa, a monolingual contextualized language representation model

Milan Straka, Jakub Náplava, Jana Straková +1

We present RobeCzech, a monolingual RoBERTa language representation model trained on Czech data. RoBERTa is a robustly optimized Transformer-based pretraining approach. We show tha…

cs.CL20202 cited

ÚFAL at MRP 2020: Permutation-invariant Semantic Parsing in PERIN

David Samuel, Milan Straka

We present PERIN, a novel permutation-invariant approach to sentence-to-graph semantic parsing. PERIN is a versatile, cross-framework and language independent architecture for univ…

cs.CL2020

Reading Comprehension in Czech via Machine Translation and Cross-lingual Transfer

Kateřina Macková, Milan Straka

Reading comprehension is a well studied task, with huge training datasets in English. This work focuses on building reading comprehension systems for Czech, without requiring any m…