works on

From the 2 of 5 linked papers with an AI index.

collaborators

5 papers

cs.CY2026

Probabilistic "Copies" in Generative AI Models

Mark A. Lemley, A. Feder Cooper

The paper examines how large language models may probabilistically reproduce copyrighted text they have memorized, and argues that current copyright law will likely treat such mode…

cs.LG2026

Extractable Memorization From First Principles

A. Feder Cooper, Marika Swanberg, Jamie Hayes +5

The paper introduces formal matched‑comparison methods—using conformal testing and document‑level censuses—to reliably determine when a language model has memorized training data,…

cs.CL2026

Estimating near-verbatim extraction risk in language models with decoding-constrained beam search

A. Feder Cooper, Mark A. Lemley, Christopher De Sa +6

Recent work shows that standard greedy-decoding extraction methods for quantifying memorization in LLMs miss how extraction risk varies across sequences. Probabilistic extraction -…

cs.CR2026

Exploring the limits of strong membership inference attacks on large language models

Jamie Hayes, Ilia Shumailov, Christopher A. Choquette-Choo +13

State-of-the-art membership inference attacks (MIAs) typically require training many reference models, making it difficult to scale these attacks to large pre-trained language mode…

cs.CY2025

The Files are in the Computer: On Copyright, Memorization, and Generative AI

A. Feder Cooper, James Grimmelmann

The New York Times's copyright lawsuit against OpenAI and Microsoft alleges OpenAI's GPT models have "memorized" NYT articles. Other lawsuits make similar claims. But parties, cour…