2 citations · 2 across the 1 of their papers we have counts for
3 papers
Extracting memorized pieces of (copyrighted) books from open-weight language models
A. Feder Cooper, Mark A. Lemley, Allison Casasola +6
Plaintiffs and defendants in copyright lawsuits over generative AI often make sweeping, opposing claims about the extent to which large language models (LLMs) memorize protected ex…
Estimating near-verbatim extraction risk in language models with decoding-constrained beam search
A. Feder Cooper, Mark A. Lemley, Christopher De Sa +6
Recent work shows that standard greedy-decoding extraction methods for quantifying memorization in LLMs miss how extraction risk varies across sequences. Probabilistic extraction -…
Beyond Scale: The Diversity Coefficient as a Data Quality Metric for Variability in Natural Language Data
Brando Miranda, Alycia Lee, Sudharsan Sundar +4
Current trends in pre-training Large Language Models (LLMs) primarily focus on the scaling of model and dataset size. While the quality of pre-training data is considered an import…