5 citations · 7 across the 7 of their papers we have counts for
7 papers
Different Tokenization Schemes Lead to Comparable Performance in Spanish Number Agreement
Catherine Arnett, Pamela D. Rivière, Tyler A. Chang +1
The relationship between language model tokenization and performance is an open area of research. Here, we investigate how different tokenization schemes impact number agreement in…
Detecting Hallucination and Coverage Errors in Retrieval Augmented Generation for Controversial Topics
Tyler A. Chang, Katrin Tomanek, Jessica Hoffmann +4
We explore a strategy to handle controversial topics in LLM-based chatbots based on Wikipedia's Neutral Point of View (NPOV) principle: acknowledge the absence of a single true ans…
A Bit of a Problem: Measurement Disparities in Dataset Sizes Across Languages
Catherine Arnett, Tyler A. Chang, Benjamin K. Bergen
How should text dataset sizes be compared across languages? Even for content-matched (parallel) corpora, UTF-8 encoded text can require a dramatically different number of bytes for…
When Is Multilinguality a Curse? Language Modeling for 250 High- and Low-Resource Languages
Tyler A. Chang, Catherine Arnett, Zhuowen Tu +1
Multilingual language models are widely used to extend NLP systems to low-resource languages. However, concrete evidence for the effects of multilinguality on language modeling per…
Structural Priming Demonstrates Abstract Grammatical Representations in Multilingual Language Models
James A. Michaelov, Catherine Arnett, Tyler A. Chang +1
Abstract grammatical knowledge - of parts of speech and grammatical patterns - is key to the capacity for linguistic generalization in humans. But how abstract is grammatical knowl…
Crosslingual Structural Priming and the Pre-Training Dynamics of Bilingual Language Models
Catherine Arnett, Tyler A. Chang, James A. Michaelov +1
Do multilingual language models share abstract grammatical representations across languages, and if so, when do these develop? Following Sinclair et al. (2022), we use structural p…