3 papers
cs.CL2026
TokSuite: Measuring the Impact of Tokenizer Choice on Language Model Behavior
Gül Sena AltıntaÅ, Malikeh Ehghaghi, Brian Lester +4
Tokenizers provide the fundamental basis through which text is represented and processed by language models (LMs). Despite the importance of tokenization, its role in LM performanc…
cs.CL2025
The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text
Nikhil Kandpal, Brian Lester, Colin Raffel +24
Large language models (LLMs) are typically trained on enormous quantities of unlicensed text, a practice that has led to scrutiny due to possible intellectual property infringement…
cs.CL2024
Training LLMs over Neurally Compressed Text
Brian Lester, Jaehoon Lee, Alex Alemi +4
In this paper, we explore the idea of training large language models (LLMs) over highly compressed text. While standard subword tokenizers compress text by a small factor, neural t…