activity
20242026
collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2026

Goldfish: Monolingual Language Models for 350 Languages

Tyler A. Chang, Catherine Arnett, Zhuowen Tu +1

For many low-resource languages, the only available language models are large multilingual models trained on many languages simultaneously. Despite state-of-the-art performance on…

cs.CL2025

Bigram Subnetworks: Mapping to Next Tokens in Transformer Language Models

Tyler A. Chang, Benjamin K. Bergen

In Transformer language models, activation vectors transform from current token embeddings to next token predictions as they pass through the model. To isolate a minimal form of th…

cs.CL2025

On the Acquisition of Shared Grammatical Representations in Bilingual Language Models

Catherine Arnett, Tyler A. Chang, James A. Michaelov +1

Crosslingual transfer is crucial to contemporary language models' multilingual capabilities, but how it occurs is not well understood. We ask what happens to a monolingual language…

cs.CL2024

Characterizing Learning Curves During Language Model Pre-Training: Learning, Forgetting, and Stability

Tyler A. Chang, Zhuowen Tu, Benjamin K. Bergen

How do language models learn to make predictions during pre-training? To study this, we extract learning curves from five autoregressive English language model pre-training runs, f…

cs.CL2024

Different Tokenization Schemes Lead to Comparable Performance in Spanish Number Agreement

Catherine Arnett, Pamela D. Rivière, Tyler A. Chang +1

The relationship between language model tokenization and performance is an open area of research. Here, we investigate how different tokenization schemes impact number agreement in…

cs.CL2024

A Bit of a Problem: Measurement Disparities in Dataset Sizes Across Languages

Catherine Arnett, Tyler A. Chang, Benjamin K. Bergen

How should text dataset sizes be compared across languages? Even for content-matched (parallel) corpora, UTF-8 encoded text can require a dramatically different number of bytes for…