◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Catherine Arnett

4 papers hereh-index 459 citations5 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author1
  • middle author2
  • last author1

Across the 4 of 4 papers where every author was matched, so the position is known.

fields
  • cs.CL4
same name
  • Catherine Arnett — 11 papers, h 7

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

activity
20242026
most citedCommon Corpus: The Largest Collection of Ethical Data for LLM Pre-Training

1 citations · 1 across the 2 of their papers we have counts for

collaborators

4 papers

cs.CL2026

Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs

Guijin Son, Seungone Kim, Catherine Arnett +73

Following the recent achievement of gold-medal performance on the IMO by frontier LLMs, the community is searching for the next meaningful and challenging target for measuring LLM…

cs.CL2026★ 1 cited

Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training

Pierre-Carl Langlais, Pavel Chizhov, Catherine Arnett +7

Large Language Models (LLMs) are pre-trained on large amounts of data from different sources and domains. Such datasets often contain trillions of tokens, including large portions…

cs.CL2025

BPE Stays on SCRIPT: Structured Encoding for Robust Multilingual Pretokenization

Sander Land, Catherine Arnett

Byte Pair Encoding (BPE) tokenizers, widely used in Large Language Models, face challenges in multilingual settings, including penalization of non-Western scripts and the creation…

cs.CL2024

Toxicity of the Commons: Curating Open-Source Pre-Training Data

Catherine Arnett, Eliot Jones, Ivan P. Yamshchikov +1

Open-source large language models are becoming increasingly available and popular among researchers and practitioners. While significant progress has been made on open-weight model…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.