◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Catherine Arnett

3 papers hereh-index 459 citations5 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author1
  • middle author1
  • last author1

Across the 3 of 3 papers where every author was matched, so the position is known.

fields
  • cs.CL3
same name
  • Catherine Arnett — 5 papers
  • Catherine Arnett — 5 papers

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

most citedToxicity of the Commons: Curating Open-Source Pre-Training Data

1 citations · 1 across the 3 of their papers we have counts for

collaborators

3 papers

cs.CL2025

BPE Stays on SCRIPT: Structured Encoding for Robust Multilingual Pretokenization

Sander Land, Catherine Arnett

Byte Pair Encoding (BPE) tokenizers, widely used in Large Language Models, face challenges in multilingual settings, including penalization of non-Western scripts and the creation…

cs.CL2024★ 1 cited

Toxicity of the Commons: Curating Open-Source Pre-Training Data

Catherine Arnett, Eliot Jones, Ivan P. Yamshchikov +1

Open-source large language models are becoming increasingly available and popular among researchers and practitioners. While significant progress has been made on open-weight model…

cs.CL2024

BPE Gets Picky: Efficient Vocabulary Refinement During Tokenizer Training

Pavel Chizhov, Catherine Arnett, Elizaveta Korotkova +1

Language models can largely benefit from efficient tokenization. However, they still mostly utilize the classical BPE algorithm, a simple and reliable method. This has been shown t…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.