activity
20202022
most citedIntegrating Unsupervised Data Generation into Self-Supervised Neural Machine Translation for Low-Resource Languages

2 citations · 5 across the 6 of their papers we have counts for

collaborators

10 papers

cs.CL20221 cited

StereoKG: Data-Driven Knowledge Graph Construction for Cultural Knowledge and Stereotypes

Awantee Deshpande, Dana Ruiter, Marius Mosbach +1

Analyzing ethnic or religious bias is important for improving fairness, accountability, and transparency of natural language processing models. However, many techniques rely on hum…

cs.CL20221 cited

Exploiting Social Media Content for Self-Supervised Style Transfer

Dana Ruiter, Thomas Kleinbauer, Cristina España-Bonet +2

Recent research on style transfer takes inspiration from unsupervised neural machine translation (UNMT), learning from large amounts of non-parallel data by exploiting cycle consis…

cs.CL2022

Placing M-Phasis on the Plurality of Hate: A Feature-Based Corpus of Hate Online

Dana Ruiter, Liane Reiners, Ashwin Geet D'Sa +6

Even though hate speech (HS) online has been an important object of research in the last decade, most HS-related corpora over-simplify the phenomenon of hate by attempting to label…

cs.CL20211 cited

EdinSaar@WMT21: North-Germanic Low-Resource Multilingual NMT

Svetlana Tchistiakova, Jesujoba Alabi, Koel Dutta Chowdhury +2

We describe the EdinSaar submission to the shared task of Multilingual Low-Resource Translation for North Germanic Languages at the Sixth Conference on Machine Translation (WMT2021…

cs.CL20212 cited

Integrating Unsupervised Data Generation into Self-Supervised Neural Machine Translation for Low-Resource Languages

Dana Ruiter, Dietrich Klakow, Josef van Genabith +1

For most language combinations, parallel data is either scarce or simply unavailable. To address this, unsupervised machine translation (UMT) exploits large amounts of monolingual…

cs.CL2021

Modeling Profanity and Hate Speech in Social Media with Semantic Subspaces

Vanessa Hahn, Dana Ruiter, Thomas Kleinbauer +1

Hate speech and profanity detection suffer from data sparsity, especially for languages other than English, due to the subjective nature of the tasks and the resulting annotation i…