7 papers
emb-diversity: A Tool for Embedding-Based Measurement of Data Diversity
Cantao Su, Menan Velayuthan, Esther Ploeger +2
There is growing evidence that data diversity is crucial for developing fair and robust NLP models. However, current approaches to measure diversity remain inconsistent and fragmen…
Social Perceptions of English Spelling Variation on Twitter: A Comparative Analysis of Human and LLM Responses
Dong Nguyen, Laura Rosseel
Spelling variation (e.g. funnnn vs. fun) can influence the social perception of texts and their writers: we often have various associations with different forms of writing (is the…
We Need to Measure Data Diversity in NLP -- Better and Broader
Dong Nguyen, Esther Ploeger
Although diversity in NLP datasets has received growing attention, the question of how to measure it remains largely underexplored. This opinion paper examines the conceptual and m…
Disentangling the Roles of Representation and Selection in Data Pruning
Yupei Du, Yingjin Song, Hugh Mee Wong +3
Data pruning, selecting small but impactful subsets, offers a promising way to efficiently scale NLP model training. However, existing methods often involve many different design c…
Tokenization is Sensitive to Language Variation
Anna Wegmann, Dong Nguyen, David Jurgens
Variation in language is ubiquitous and often systematically linked to regional, social, and contextual factors. Tokenizers split texts into smaller units and might behave differen…
Identity Emergence in the Context of Vaccine Criticism in France
Melody Sepahpour-Fard, Michael Quayle, Padraig MacCarron +2
This study investigates the emergence of collective identity among individuals critical of vaccination policies in France during the COVID-19 pandemic. As concerns grew over mandat…