1 citations · 1 across the 5 of their papers we have counts for
3 papers · 1 filter
Institutional Books - Enriched Text: A customizable multilingual open-source pipeline for denoising, deduplicating, and annotating OCR text at scale
David Lowry-Duda, Matteo Cargnelutti, Catherine Brobston +4
Released in 2025, Institutional Books: Harvard Library (IB-HL) is a collection of 983,004 volumes (242B o200k_base tokens), originally digitized through Harvard Library's participa…
Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers
Matteo Cargnelutti, Catherine Brobston, Eben English +6
Historical newspapers are an abundant record of public life, but their dense, irregular and sometimes noisy layouts make computational access to these materials both challenging an…
Institutional Books 1.0: A 242B token dataset from Harvard Library's collections, refined for accuracy and usability
Matteo Cargnelutti, Catherine Brobston, John Hess +8
Large language models (LLMs) use data to learn about the world in order to produce meaningful correlations and predictions. As such, the nature, scale, quality, and diversity of th…