9 citations · 11 across the 7 of their papers we have counts for
9 papers · 1 filter
Detecting Depression in Thai Blog Posts: a Dataset and a Baseline
Mika Hämäläinen, Pattama Patpong, Khalid Alnajjar +2
We present the first openly available corpus for detecting depression in Thai. Our corpus is compiled by expert verified cases of depression in several online blogs. We experiment…
Finnish Dialect Identification: The Effect of Audio and Text
Mika Hämäläinen, Khalid Alnajjar, Niko Partanen +1
Finnish is a language with multiple dialects that not only differ from each other in terms of accent (pronunciation) but also in terms of morphological forms and lexical choice. We…
Apurinã Universal Dependencies Treebank
Jack Rueter, Marília Fernanda Pereira de Freitas, Sidney da Silva Facundes +2
This paper presents and discusses the first Universal Dependencies treebank for the Apurinã language. The treebank contains 76 fully annotated sentences, applies 14 parts-of-speech…
Never guess what I heard... Rumor Detection in Finnish News: a Dataset and a Baseline
Mika Hämäläinen, Khalid Alnajjar, Niko Partanen +1
This study presents a new dataset on rumor detection in Finnish language news headlines. We have evaluated two different LSTM based models and two different BERT models, and have f…
Neural Morphology Dataset and Models for Multiple Languages, from the Large to the Endangered
Mika Hämäläinen, Niko Partanen, Jack Rueter +1
We train neural models for morphological analysis, generation and lemmatization for morphologically rich languages. We present a method for automatically extracting substantially l…
Ve'rdd. Narrowing the Gap between Paper Dictionaries, Low-Resource NLP and Community Involvement
Khalid Alnajjar, Mika Hämäläinen, Jack Rueter +1
We present an open-source online dictionary editing system, Ve'rdd, that offers a chance to re-evaluate and edit grassroots dictionaries that have been exposed to multiple amateur…