activity
20242026
collaborators

7 papers

cs.CL2026

emb-diversity: A Tool for Embedding-Based Measurement of Data Diversity

Cantao Su, Menan Velayuthan, Esther Ploeger +2

There is growing evidence that data diversity is crucial for developing fair and robust NLP models. However, current approaches to measure diversity remain inconsistent and fragmen…

cs.CL2026

How Good is Your Wikipedia? Auditing Data Quality for Low-resource and Multilingual NLP

Kushal Tatariya, Artur Kulmizev, Wessel Poelman +6

Wikipedia's perceived high quality and broad language coverage have established it as a fundamental resource in NLP. However, in recent years, such assumptions of high quality have…

cs.CL2025

We Need to Measure Data Diversity in NLP -- Better and Broader

Dong Nguyen, Esther Ploeger

Although diversity in NLP datasets has received growing attention, the question of how to measure it remains largely underexplored. This opinion paper examines the conceptual and m…

cs.CL2025

A Principled Framework for Evaluating on Typologically Diverse Languages

Esther Ploeger, Wessel Poelman, Andreas Holck Høeg-Petersen +3

Beyond individual languages, multilingual natural language processing (NLP) research increasingly aims to develop models that perform well across languages generally. However, eval…

cs.CL2025

Multi-perspective Alignment for Increasing Naturalness in Neural Machine Translation

Huiyuan Lai, Esther Ploeger, Rik van Noord +1

Neural machine translation (NMT) systems amplify lexical biases present in their training data, leading to artificially impoverished language in output translations. These language…

cs.CL2024

INCLUDE: Evaluating Multilingual Language Understanding with Regional Knowledge

Angelika Romanou, Negar Foroutan, Anna Sotnikova +56

The performance differential of large language models (LLM) between languages hinders their effective deployment in many regions, inhibiting the potential economic and societal val…