activity
20172020
most citedEvaluating Sentence Segmentation and Word Tokenization Systems on Estonian Web Texts

11 citations · 11 across the 1 of their papers we have counts for

collaborators

7 papers

cs.CL202011 cited

Evaluating Sentence Segmentation and Word Tokenization Systems on Estonian Web Texts

Kairit Sirts, Kairit Peekman

Texts obtained from web are noisy and do not necessarily follow the orthographic sentence and word boundary rules. Thus, sentence segmentation and word tokenization systems that ha…

cs.CL2020

EstBERT: A Pretrained Language-Specific BERT for Estonian

Hasan Tanvir, Claudia Kittask, Sandra Eiche +1

This paper presents EstBERT, a large pretrained transformer-based language-specific BERT model for Estonian. Recent work has evaluated multilingual BERT models on Estonian tasks an…

cs.CL2020

Evaluating Multilingual BERT for Estonian

Claudia Kittask, Kirill Milintsevich, Kairit Sirts

Recently, large pre-trained language models, such as BERT, have reached state-of-the-art performance in many natural language processing tasks, but for many languages, including Es…

cs.CL2018

Modeling Composite Labels for Neural Morphological Tagging

Alexander Tkachenko, Kairit Sirts

Neural morphological tagging has been regarded as an extension to POS tagging task, treating each morphological tag as a monolithic label and ignoring its internal structure. We pr…

cs.CL2018

Neural Morphological Tagging for Estonian

Alexander Tkachenko, Kairit Sirts

We develop neural morphological tagging and disambiguation models for Estonian. First, we experiment with two neural architectures for morphological tagging - a standard multiclass…

cs.IR2018

The Impact of Annotation Guidelines and Annotated Data on Extracting App Features from App Reviews

Faiz Ali Shah, Kairit Sirts, Dietmar Pfahl

Annotation guidelines used to guide the annotation of training and evaluation datasets can have a considerable impact on the quality of machine learning models. In this study, we e…