5 citations · 8 across the 4 of their papers we have counts for
7 papers · 1 filter
Lies, Damned Lies, and Distributional Language Statistics: Persuasion and Deception with Large Language Models
Cameron R. Jones, Benjamin K. Bergen
Large Language Models (LLMs) can generate content that is as persuasive as human-written text and appear capable of selectively producing deceptive outputs. These capabilities rais…
Why do language models perform worse for morphologically complex languages?
Catherine Arnett, Benjamin K. Bergen
Language models perform differently across languages. It has been previously suggested that morphological typology may explain some of this variability (Cotterell et al., 2018). We…
A Bit of a Problem: Measurement Disparities in Dataset Sizes Across Languages
Catherine Arnett, Tyler A. Chang, Benjamin K. Bergen
How should text dataset sizes be compared across languages? Even for content-matched (parallel) corpora, UTF-8 encoded text can require a dramatically different number of bytes for…
When Is Multilinguality a Curse? Language Modeling for 250 High- and Low-Resource Languages
Tyler A. Chang, Catherine Arnett, Zhuowen Tu +1
Multilingual language models are widely used to extend NLP systems to low-resource languages. However, concrete evidence for the effects of multilinguality on language modeling per…
Structural Priming Demonstrates Abstract Grammatical Representations in Multilingual Language Models
James A. Michaelov, Catherine Arnett, Tyler A. Chang +1
Abstract grammatical knowledge - of parts of speech and grammatical patterns - is key to the capacity for linguistic generalization in humans. But how abstract is grammatical knowl…
Crosslingual Structural Priming and the Pre-Training Dynamics of Bilingual Language Models
Catherine Arnett, Tyler A. Chang, James A. Michaelov +1
Do multilingual language models share abstract grammatical representations across languages, and if so, when do these develop? Following Sinclair et al. (2022), we use structural p…