11 papers
QQ: A Language Metadata Toolkit for Multilingual NLP
Wessel Poelman, Yiyi Chen, Miryam de Lhoneux
Multilingual NLP research increasingly involves hundreds or thousands of languages across different datasets. Managing, discovering, and reporting language metadata becomes a commo…
How Good is Your Wikipedia? Auditing Data Quality for Low-resource and Multilingual NLP
Kushal Tatariya, Artur Kulmizev, Wessel Poelman +6
Wikipedia's perceived high quality and broad language coverage have established it as a fundamental resource in NLP. However, in recent years, such assumptions of high quality have…
Typologically Informed Parameter Aggregation
Stef Accou, Wessel Poelman
Massively multilingual language models enable cross-lingual generalization but underperform on low-resource and unseen languages. While adapter-based fine-tuning offers a parameter…
Form and Meaning in Intrinsic Multilingual Evaluations
Wessel Poelman, Miryam de Lhoneux
Intrinsic evaluation metrics for conditional language models, such as perplexity or bits-per-character, are widely used in both mono- and multilingual settings. These metrics are r…
On the Interplay between Positional Encodings, Morphological Complexity, and Word Order Flexibility
Kushal Tatariya, Wessel Poelman, Miryam de Lhoneux
Language model architectures are predominantly first created for English and subsequently applied to other languages. It is an open question whether this architectural bias leads t…
Confounding Factors in Relating Model Performance to Morphology
Wessel Poelman, Thomas Bauwens, Miryam de Lhoneux
The extent to which individual language characteristics influence tokenization and language modeling is an open question. Differences in morphological systems have been suggested a…