3 papers
cs.DL2025
GRIN Transfer: A production-ready tool for libraries to retrieve digital copies from Google Books
Liza Daly, Matteo Cargnelutti, Catherine Brobston +4
Publicly launched in 2004, the Google Books project has scanned tens of millions of items in partnership with libraries around the world. As part of this project, Google created th…
cs.CL2025
Institutional Books 1.0: A 242B token dataset from Harvard Library's collections, refined for accuracy and usability
Matteo Cargnelutti, Catherine Brobston, John Hess +8
Large language models (LLMs) use data to learn about the world in order to produce meaningful correlations and predictions. As such, the nature, scale, quality, and diversity of th…
cs.LG2024
SEAL: Systematic Error Analysis for Value ALignment
Manon Revel, Matteo Cargnelutti, Tyna Eloundou +1
Reinforcement Learning from Human Feedback (RLHF) aims to align language models (LMs) with human values by training reward models (RMs) on binary preferences and using these RMs to…