◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Brian Lester

Google

3 papers hereh-index 1211.6k citations21 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author1
  • middle author2

Across the 3 of 3 papers where every author was matched, so the position is known.

fields
  • cs.CL3
affiliations
  • Google
Homepage
same name
  • Brian Lester — 1 paper, h 1
  • Brian Lester — 1 paper, h 1

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

activity
20242026
collaborators

3 papers

cs.CL2026

TokSuite: Measuring the Impact of Tokenizer Choice on Language Model Behavior

Gül Sena Altıntaş, Malikeh Ehghaghi, Brian Lester +4

Tokenizers provide the fundamental basis through which text is represented and processed by language models (LMs). Despite the importance of tokenization, its role in LM performanc…

cs.CL2025

The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text

Nikhil Kandpal, Brian Lester, Colin Raffel +24

Large language models (LLMs) are typically trained on enormous quantities of unlicensed text, a practice that has led to scrutiny due to possible intellectual property infringement…

cs.CL2024

Training LLMs over Neurally Compressed Text

Brian Lester, Jaehoon Lee, Alex Alemi +4

In this paper, we explore the idea of training large language models (LLMs) over highly compressed text. While standard subword tokenizers compress text by a small factor, neural t…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.