activity
20172022
most citedNeural machine translation for low-resource languages

30 citations · 56 across the 10 of their papers we have counts for

collaborators

20 papers

cs.CL20222 cited

When to Laugh and How Hard? A Multimodal Approach to Detecting Humor and its Intensity

Khalid Alnajjar, Mika Hämäläinen, Jörg Tiedemann +2

Prerecorded laughter accompanying dialog in comedy TV shows encourages the audience to laugh by clearly marking humorous moments in the show. We present an approach for automatical…

cs.CL202115 cited

Analyzing the Use of Character-Level Translation with Sparse and Noisy Datasets

Jörg Tiedemann, Preslav Nakov

This paper provides an analysis of character-level machine translation models used in pivot-based translation when applied to sparse and noisy datasets, such as crowdsourced movie…

cs.CL20212 cited

NLI Data Sanity Check: Assessing the Effect of Data Corruption on Model Performance

Aarne Talman, Marianna Apidianaki, Stergios Chatzikyriakidis +1

Pre-trained neural language models give high performance on natural language inference (NLI) tasks. But whether they actually understand the meaning of the processed sequences rema…

cs.CL20206 cited

XED: A Multilingual Dataset for Sentiment Analysis and Emotion Detection

Emily Öhman, Marc Pàmies, Kaisla Kajava +1

We introduce XED, a multilingual fine-grained emotion dataset. The dataset consists of human-annotated Finnish (25k) and English sentences (30k), as well as projected annotations f…

cs.CL2020

The Tatoeba Translation Challenge -- Realistic Data Sets for Low Resource and Multilingual MT

Jörg Tiedemann

This paper describes the development of a new benchmark for machine translation that provides training and test data for thousands of language pairs covering over 500 languages and…

cs.CL2020

LT@Helsinki at SemEval-2020 Task 12: Multilingual or language-specific BERT?

Marc Pàmies, Emily Öhman, Kaisla Kajava +1

This paper presents the different models submitted by the LT@Helsinki team for the SemEval 2020 Shared Task 12. Our team participated in sub-tasks A and C; titled offensive languag…