most citedVe'rdd. Narrowing the Gap between Paper Dictionaries, Low-Resource NLP and Community Involvement

9 citations · 20 across the 11 of their papers we have counts for

collaborators

13 papers

cs.CL2021

Detecting Depression in Thai Blog Posts: a Dataset and a Baseline

Mika Hämäläinen, Pattama Patpong, Khalid Alnajjar +2

We present the first openly available corpus for detecting depression in Thai. Our corpus is compiled by expert verified cases of depression in several online blogs. We experiment…

cs.CL20211 cited

Finnish Dialect Identification: The Effect of Audio and Text

Mika Hämäläinen, Khalid Alnajjar, Niko Partanen +1

Finnish is a language with multiple dialects that not only differ from each other in terms of accent (pronunciation) but also in terms of morphological forms and lexical choice. We…

cs.CL2021

How Cute is Pikachu? Gathering and Ranking Pokémon Properties from Data with Pokémon Word Embeddings

Mika Hämäläinen, Khalid Alnajjar, Niko Partanen

We present different methods for obtaining descriptive properties automatically for the 151 original Pokémon. We train several different word embeddings models on a crawled Pokémon…

cs.CL20214 cited

Lemmatization of Historical Old Literary Finnish Texts in Modern Orthography

Mika Hämäläinen, Niko Partanen, Khalid Alnajjar

Texts written in Old Literary Finnish represent the first literary work ever written in Finnish starting from the 16th century. There have been several projects in Finland that hav…

cs.CL2021

Apurinã Universal Dependencies Treebank

Jack Rueter, Marília Fernanda Pereira de Freitas, Sidney da Silva Facundes +2

This paper presents and discusses the first Universal Dependencies treebank for the Apurinã language. The treebank contains 76 fully annotated sentences, applies 14 parts-of-speech…

cs.CL2021

Never guess what I heard... Rumor Detection in Finnish News: a Dataset and a Baseline

Mika Hämäläinen, Khalid Alnajjar, Niko Partanen +1

This study presents a new dataset on rumor detection in Finnish language news headlines. We have evaluated two different LSTM based models and two different BERT models, and have f…