2 papers
cs.CY2025
Towards Best Practices for Open Datasets for LLM Training
Stefan Baack, Stella Biderman, Kasia Odrozek +36
Many AI companies are training their large language models (LLMs) on data without the permission of the copyright owners. The permissibility of doing so varies by jurisdiction: in…
cs.CY2024
My Voice, Your Voice, Our Voice: Attitudes Towards Collective Governance of a Choral AI Dataset
Jennifer Ding, Eva Jäger, Victoria Ivanova +1
Data grows in value when joined and combined; likewise the power of voice grows in ensemble. With 15 UK choirs, we explore opportunities for bottom-up data governance of a jointly…