2 papers
cs.CL2025
Institutional Books 1.0: A 242B token dataset from Harvard Library's collections, refined for accuracy and usability
Matteo Cargnelutti, Catherine Brobston, John Hess +8
Large language models (LLMs) use data to learn about the world in order to produce meaningful correlations and predictions. As such, the nature, scale, quality, and diversity of th…
cs.HC2025
Creative Writers' Attitudes on Writing as Training Data for Large Language Models
Katy Ilonka Gero, Meera Desai, Carly Schnitzler +3
The use of creative writing as training data for large language models (LLMs) is highly contentious and many writers have expressed outrage at the use of their work without consent…