activity
20242026
collaborators

7 papers

cs.CL2026

Estonian WinoGrande Dataset: Comparative Analysis of LLM Performance on Human and Machine Translation

Marii Ojastu, Hele-Andra Kuulmets, Aleksei Dorkin +3

In this paper, we present a localized and culturally adapted Estonian translation of the test set from the widely used commonsense reasoning benchmark, WinoGrande. We detail the tr…

cs.CL2026

EstLLM: Enhancing Estonian Capabilities in Multilingual LLMs via Continued Pretraining and Post-Training

Aleksei Dorkin, Taido Purason, Emil Kalbaliyev +7

Large language models (LLMs) are predominantly trained on English-centric data, resulting in uneven performance for smaller languages. We study whether continued pretraining (CPT)…

cs.CL2025

TartuNLP at SemEval-2025 Task 5: Subject Tagging as Two-Stage Information Retrieval

Aleksei Dorkin, Kairit Sirts

We present our submission to the Task 5 of SemEval-2025 that aims to aid librarians in assigning subject tags to the library records by producing a list of likely relevant tags for…

cs.CL2025

GliLem: Leveraging GliNER for Contextualized Lemmatization in Estonian

Aleksei Dorkin, Kairit Sirts

We present GliLem -- a novel hybrid lemmatization system for Estonian that enhances the highly accurate rule-based morphological analyzer Vabamorf with an external disambiguation m…

cs.CL2025

Prune or Retrain: Optimizing the Vocabulary of Multilingual Models for Estonian

Aleksei Dorkin, Taido Purason, Kairit Sirts

Adapting multilingual language models to specific languages can enhance both their efficiency and performance. In this study, we explore how modifying the vocabulary of a multiling…

cs.CL2024

TartuNLP at EvaLatin 2024: Emotion Polarity Detection

Aleksei Dorkin, Kairit Sirts

This paper presents the TartuNLP team submission to EvaLatin 2024 shared task of the emotion polarity detection for historical Latin texts. Our system relies on two distinct approa…