11 citations · 20 across the 6 of their papers we have counts for
16 papers
Quality and Efficiency of Manual Annotation: Pre-annotation Bias
Marie Mikulová, Milan Straka, Jan Štěpánek +2
This paper presents an analysis of annotation using an automatic pre-annotation for a mid-level annotation complexity task -- dependency syntax annotation. It compares the annotati…
DaMuEL: A Large Multilingual Dataset for Entity Linking
David Kubeša, Milan Straka
We present DaMuEL, a large Multilingual Dataset for Entity Linking containing data in 53 languages. DaMuEL consists of two components: a knowledge base that contains language-agnos…
Diacritics Restoration using BERT with Analysis on Czech language
Jakub Náplava, Milan Straka, Jana Straková
We propose a new architecture for diacritics restoration based on contextualized embeddings, namely BERT, and we evaluate it on 12 languages with diacritics. Furthermore, we conduc…
RobeCzech: Czech RoBERTa, a monolingual contextualized language representation model
Milan Straka, Jakub Náplava, Jana Straková +1
We present RobeCzech, a monolingual RoBERTa language representation model trained on Czech data. RoBERTa is a robustly optimized Transformer-based pretraining approach. We show tha…
ÚFAL at MRP 2020: Permutation-invariant Semantic Parsing in PERIN
David Samuel, Milan Straka
We present PERIN, a novel permutation-invariant approach to sentence-to-graph semantic parsing. PERIN is a versatile, cross-framework and language independent architecture for univ…
Reading Comprehension in Czech via Machine Translation and Cross-lingual Transfer
Kateřina Macková, Milan Straka
Reading comprehension is a well studied task, with huge training datasets in English. This work focuses on building reading comprehension systems for Czech, without requiring any m…