activity
20242026
collaborators

5 papers

cs.CL2026

How Good is Your Wikipedia? Auditing Data Quality for Low-resource and Multilingual NLP

Kushal Tatariya, Artur Kulmizev, Wessel Poelman +6

Wikipedia's perceived high quality and broad language coverage have established it as a fundamental resource in NLP. However, in recent years, such assumptions of high quality have…

cs.CL2025

Beyond Weaponization: NLP Security for Medium and Lower-Resourced Languages in Their Own Right

Heather Lent

Despite mounting evidence that multilinguality can be easily weaponized against language models (LMs), works across NLP Security remain overwhelmingly English-centric. In terms of…

cs.CL2025

Limited-Resource Adapters Are Regularizers, Not Linguists

Marcell Fekete, Nathaniel R. Robinson, Ernests Lavrinovics +4

Cross-lingual transfer from related high-resource languages is a well-established strategy to enhance low-resource language technologies. Prior work has shown that adapters show pr…

cs.CL2025

NLP Security and Ethics, in the Wild

Heather Lent, Erick Galinkin, Yiyi Chen +3

As NLP models are used by a growing number of end-users, an area of increasing importance is NLP Security (NLPSec): assessing the vulnerability of models to malicious attacks and d…

cs.CL2024

Against All Odds: Overcoming Typology, Script, and Language Confusion in Multilingual Embedding Inversion Attacks

Yiyi Chen, Russa Biswas, Heather Lent +1

Large Language Models (LLMs) are susceptible to malicious influence by cyber attackers through intrusions such as adversarial, backdoor, and embedding inversion attacks. In respons…