10 papers
Left-Branching Transformers Excel at Right-Branching Languages: Data Shapes Word Order Preferences in Language Models
Varvara Arzt, Allan Hanbury, Terra Blevins
We systematically compare word order preferences in decoder-only language models across 192 artificial languages and typologically diverse natural languages. On artificial language…
Parameter Alignment Mitigates Catastrophic Forgetting in Multilingual Expert Language Models
Sanchit Ahuja, Terra Blevins
While continual pretraining~(CPT) is a practical way to extend large language models to new languages, naïve finetuning on targeted data erodes existing capabilities through catas…
Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
Tyler A. Chang, Catherine Arnett, Abdelrahman Sadallah +377
To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we pre…
Accommodation Goes Both Ways: Studying Linguistic Convergence Between Humans and Language Models
Terra Blevins
As LLMs become increasingly integrated into daily life, understanding how their presence will shape human linguistic behavior is an open question. We present a large-scale study of…
Calibration Is Not Enough: Evaluating Confidence Estimation Under Language Variations
Yuxi Xia, Dennis Ulmer, Terra Blevins +3
Confidence estimation (CE) indicates how reliable the answers of large language models are and impacts user trust and decision-making. Existing evaluations mainly concern the align…
Universal NER v2: Towards a Massively Multilingual Named Entity Recognition Benchmark
Terra Blevins, Stephen Mayhew, Marek Å uppa +11
While multilingual language models promise to bring the benefits of LLMs to speakers of many languages, gold-standard evaluation benchmarks in most languages to interrogate these a…