120 citations · 220 across the 16 of their papers we have counts for
5 papers · 2 filters
Multilingual is not enough: BERT for Finnish
Antti Virtanen, Jenna Kanerva, Rami Ilo +5
Deep learning-based language models pretrained on large unannotated text corpora have been demonstrated to allow efficient transfer learning for natural language processing, with r…
Morphological Tagging and Lemmatization of Albanian: A Manually Annotated Corpus and Neural Models
Nelda Kote, Marenglen Biba, Jenna Kanerva +2
In this paper, we present the first publicly available part-of-speech and morphologically tagged corpus for the Albanian language, as well as a neural morphological tagger and lemm…
Is Multilingual BERT Fluent in Language Generation?
Samuel Rönnqvist, Jenna Kanerva, Tapio Salakoski +1
The multilingual BERT model is trained on 104 languages and meant to serve as a universal language model and tool for encoding sentences. We explore how well the model performs on…
Template-free Data-to-Text Generation of Finnish Sports News
Jenna Kanerva, Samuel Rönnqvist, Riina Kekki +2
News articles such as sports game reports are often thought to closely follow the underlying game statistics, but in practice they contain a notable amount of background knowledge,…
Universal Lemmatizer: A Sequence to Sequence Model for Lemmatizing Universal Dependencies Treebanks
Jenna Kanerva, Filip Ginter, Tapio Salakoski
In this paper we present a novel lemmatization method based on a sequence-to-sequence neural network architecture and morphosyntactic context representation. In the proposed method…