From the 2 of 5 linked papers with an AI index.
5 papers
Probabilistic "Copies" in Generative AI Models
Mark A. Lemley, A. Feder Cooper
The paper examines how large language models may probabilistically reproduce copyrighted text they have memorized, and argues that current copyright law will likely treat such mode…
Extractable Memorization From First Principles
A. Feder Cooper, Marika Swanberg, Jamie Hayes +5
The paper introduces formal matched‑comparison methods—using conformal testing and document‑level censuses—to reliably determine when a language model has memorized training data,…
Estimating near-verbatim extraction risk in language models with decoding-constrained beam search
A. Feder Cooper, Mark A. Lemley, Christopher De Sa +6
Recent work shows that standard greedy-decoding extraction methods for quantifying memorization in LLMs miss how extraction risk varies across sequences. Probabilistic extraction -…
Exploring the limits of strong membership inference attacks on large language models
Jamie Hayes, Ilia Shumailov, Christopher A. Choquette-Choo +13
State-of-the-art membership inference attacks (MIAs) typically require training many reference models, making it difficult to scale these attacks to large pre-trained language mode…
The Files are in the Computer: On Copyright, Memorization, and Generative AI
A. Feder Cooper, James Grimmelmann
The New York Times's copyright lawsuit against OpenAI and Microsoft alleges OpenAI's GPT models have "memorized" NYT articles. Other lawsuits make similar claims. But parties, cour…