◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Catherine Brobston

5 papers hereh-index 15 citations6 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author5

Across the 5 of 5 papers where every author was matched, so the position is known.

fields
  • cs.CL3
  • cs.CV1
  • cs.DL1

identity via Semantic Scholar / OpenAlex

most citedInstitutional Books 1.0: A 242B token dataset from Harvard Library's collections, refined for accuracy and usability

1 citations · 1 across the 5 of their papers we have counts for

collaborators
Showing cs.CLShow all

3 papers · 1 filter

cs.CL2026

Institutional Books - Enriched Text: A customizable multilingual open-source pipeline for denoising, deduplicating, and annotating OCR text at scale

David Lowry-Duda, Matteo Cargnelutti, Catherine Brobston +4

Released in 2025, Institutional Books: Harvard Library (IB-HL) is a collection of 983,004 volumes (242B o200k_base tokens), originally digitized through Harvard Library's participa…

cs.CL2026

Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers

Matteo Cargnelutti, Catherine Brobston, Eben English +6

Historical newspapers are an abundant record of public life, but their dense, irregular and sometimes noisy layouts make computational access to these materials both challenging an…

cs.CL2025★ 1 cited

Institutional Books 1.0: A 242B token dataset from Harvard Library's collections, refined for accuracy and usability

Matteo Cargnelutti, Catherine Brobston, John Hess +8

Large language models (LLMs) use data to learn about the world in order to produce meaningful correlations and predictions. As such, the nature, scale, quality, and diversity of th…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.