most citedExtracting memorized pieces of (copyrighted) books from open-weight language models

2 citations · 2 across the 1 of their papers we have counts for

collaborators

5 papers

cs.CL20262 cited

Extracting memorized pieces of (copyrighted) books from open-weight language models

A. Feder Cooper, Mark A. Lemley, Allison Casasola +6

Plaintiffs and defendants in copyright lawsuits over generative AI often make sweeping, opposing claims about the extent to which large language models (LLMs) memorize protected ex…

cs.CL2026

Extracting books from production language models

Ahmed Ahmed, A. Feder Cooper, Sanmi Koyejo +1

Many unresolved legal questions over LLMs and copyright center on memorization: whether specific training data have been encoded in the model's weights during training, and whether…

cs.CL2025

SpecEval: Evaluating Model Adherence to Behavior Specifications

Ahmed Ahmed, Kevin Klyman, Yi Zeng +2

Companies that develop foundation models publish behavioral guidelines they pledge their models will follow, but it remains unclear if models actually do so. While providers such a…

cs.CY2025

AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons

Shaona Ghosh, Heather Frase, Adina Williams +99

The rapid advancement and deployment of AI systems have created an urgent need for standard safety-evaluation frameworks. This paper introduces AILuminate v1.0, the first comprehen…

cs.LG2025

Independence Tests for Language Models

Sally Zhu, Ahmed Ahmed, Rohith Kuditipudi +1

We consider the following problem: given the weights of two models, can we test whether they were trained independently -- i.e., from independent random initializations? We conside…