30 citations · 56 across the 10 of their papers we have counts for
20 papers
When to Laugh and How Hard? A Multimodal Approach to Detecting Humor and its Intensity
Khalid Alnajjar, Mika Hämäläinen, Jörg Tiedemann +2
Prerecorded laughter accompanying dialog in comedy TV shows encourages the audience to laugh by clearly marking humorous moments in the show. We present an approach for automatical…
Analyzing the Use of Character-Level Translation with Sparse and Noisy Datasets
Jörg Tiedemann, Preslav Nakov
This paper provides an analysis of character-level machine translation models used in pivot-based translation when applied to sparse and noisy datasets, such as crowdsourced movie…
NLI Data Sanity Check: Assessing the Effect of Data Corruption on Model Performance
Aarne Talman, Marianna Apidianaki, Stergios Chatzikyriakidis +1
Pre-trained neural language models give high performance on natural language inference (NLI) tasks. But whether they actually understand the meaning of the processed sequences rema…
XED: A Multilingual Dataset for Sentiment Analysis and Emotion Detection
Emily Öhman, Marc Pàmies, Kaisla Kajava +1
We introduce XED, a multilingual fine-grained emotion dataset. The dataset consists of human-annotated Finnish (25k) and English sentences (30k), as well as projected annotations f…
The Tatoeba Translation Challenge -- Realistic Data Sets for Low Resource and Multilingual MT
Jörg Tiedemann
This paper describes the development of a new benchmark for machine translation that provides training and test data for thousands of language pairs covering over 500 languages and…
LT@Helsinki at SemEval-2020 Task 12: Multilingual or language-specific BERT?
Marc Pàmies, Emily Öhman, Kaisla Kajava +1
This paper presents the different models submitted by the LT@Helsinki team for the SemEval 2020 Shared Task 12. Our team participated in sub-tasks A and C; titled offensive languag…