2 citations · 2 across the 1 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text
Nikhil Kandpal, Brian Lester, Colin Raffel +24
Large language models (LLMs) are typically trained on enormous quantities of unlicensed text, a practice that has led to scrutiny due to possible intellectual property infringement…
cs.CL2025★ 2 cited
Extracting memorized pieces of (copyrighted) books from open-weight language models
A. Feder Cooper, Mark A. Lemley, Allison Casasola +6
Plaintiffs and defendants in copyright lawsuits over generative AI often make sweeping, opposing claims about the extent to which large language models (LLMs) memorize protected ex…