most citedWhy do language models perform worse for morphologically complex languages?

1 citations · 1 across the 1 of their papers we have counts for

collaborators

5 papers

cs.CL2025

Disaggregation Reveals Hidden Training Dynamics: The Case of Agreement Attraction

James A. Michaelov, Catherine Arnett

Language models generally produce grammatical text, but they are more likely to make errors in certain contexts. Drawing on paradigms from psycholinguistics, we carry out a fine-gr…

cs.CL2025

Explaining and Mitigating Crosslingual Tokenizer Inequities

Catherine Arnett, Tyler A. Chang, Stella Biderman +1

The number of tokens it takes to encode parallel text in different languages is known to vary. These disparities are called token premiums. Having high token premiums leads to less…

cs.CL2025

Evaluating Morphological Alignment of Tokenizers in 70 Languages

Catherine Arnett, Marisa Hudspeth, Brendan O'Connor

While tokenization is a key step in language modeling, with effects on model training and performance, it remains unclear how to effectively evaluate tokenizer quality. One propose…

cs.CL2025

On the Acquisition of Shared Grammatical Representations in Bilingual Language Models

Catherine Arnett, Tyler A. Chang, James A. Michaelov +1

Crosslingual transfer is crucial to contemporary language models' multilingual capabilities, but how it occurs is not well understood. We ask what happens to a monolingual language…

cs.CL20241 cited

Why do language models perform worse for morphologically complex languages?

Catherine Arnett, Benjamin K. Bergen

Language models perform differently across languages. It has been previously suggested that morphological typology may explain some of this variability (Cotterell et al., 2018). We…