large language models 2conformal inference 1copyright 1extractable memorization 1legal analysis 1memorization 1privacy evaluation 1probabilistic generation 1statistical testing 1
From the 2 of 4 linked papers with an AI index.
3 papers
cs.CY2026
Probabilistic "Copies" in Generative AI Models
Mark A. Lemley, A. Feder Cooper
The paper examines how large language models may probabilistically reproduce copyrighted text they have memorized, and argues that current copyright law will likely treat such mode…
cs.LG2026
Extractable Memorization From First Principles
A. Feder Cooper, Marika Swanberg, Jamie Hayes +5
The paper introduces formal matched‑comparison methods—using conformal testing and document‑level censuses—to reliably determine when a language model has memorized training data,…
cs.CL2026
Estimating near-verbatim extraction risk in language models with decoding-constrained beam search
A. Feder Cooper, Mark A. Lemley, Christopher De Sa +6
Recent work shows that standard greedy-decoding extraction methods for quantifying memorization in LLMs miss how extraction risk varies across sequences. Probabilistic extraction -…