activity
20192026
most citedMultilingual is not enough: BERT for Finnish

120 citations · 220 across the 16 of their papers we have counts for

collaborators
Showing 2019 · cs.CLShow all

5 papers · 2 filters

cs.CL2019★ 120 cited

Multilingual is not enough: BERT for Finnish

Antti Virtanen, Jenna Kanerva, Rami Ilo +5

Deep learning-based language models pretrained on large unannotated text corpora have been demonstrated to allow efficient transfer learning for natural language processing, with r…

cs.CL2019★ 8 cited

Morphological Tagging and Lemmatization of Albanian: A Manually Annotated Corpus and Neural Models

Nelda Kote, Marenglen Biba, Jenna Kanerva +2

In this paper, we present the first publicly available part-of-speech and morphologically tagged corpus for the Albanian language, as well as a neural morphological tagger and lemm…

cs.CL2019★ 39 cited

Is Multilingual BERT Fluent in Language Generation?

Samuel Rönnqvist, Jenna Kanerva, Tapio Salakoski +1

The multilingual BERT model is trained on 104 languages and meant to serve as a universal language model and tool for encoding sentences. We explore how well the model performs on…

cs.CL2019★ 10 cited

Template-free Data-to-Text Generation of Finnish Sports News

Jenna Kanerva, Samuel Rönnqvist, Riina Kekki +2

News articles such as sports game reports are often thought to closely follow the underlying game statistics, but in practice they contain a notable amount of background knowledge,…

cs.CL2019

Universal Lemmatizer: A Sequence to Sequence Model for Lemmatizing Universal Dependencies Treebanks

Jenna Kanerva, Filip Ginter, Tapio Salakoski

In this paper we present a novel lemmatization method based on a sequence-to-sequence neural network architecture and morphosyntactic context representation. In the proposed method…