NewEvery arXiv paper, its researchers & institutions — mapped.
papers

Publications (38)

astro-ph.GA2012

Filamentary Star Formation: Observing the Evolution toward Flattened Envelopes

Katherine Lee, Leslie Looney, Doug Johnstone +1

cs.LG2025

Machine Unlearning Doesn't Do What You Think: Lessons for Generative AI Policy and Research

A. Feder Cooper, Christopher A. Choquette-Choo, Miranda Bogen +34

cs.CR2021

Extracting Training Data from Large Language Models

Nicholas Carlini, Florian Tramer, Eric Wallace +9

cs.CL2025

Gemini: A Family of Highly Capable Multimodal Models

Gemini Team, Rohan Anil, Sebastian Borgeaud +1340

cs.LG2023

Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

Colin Raffel, Noam Shazeer, Adam Roberts +6

cs.LG2023

Reverse-Engineering Decoding Strategies Given Blackbox Access to a Language Generation System

Daphne Ippolito, Nicholas Carlini, Katherine Lee +2

cs.CL2022

Deduplicating Training Data Makes Language Models Better

Katherine Lee, Daphne Ippolito, Andrew Nystrom +4

cs.LG2025

Exploring and Mitigating Adversarial Manipulation of Voting-Based Leaderboards

Yangsibo Huang, Milad Nasr, Anastasios Angelopoulos +10

cs.CL2024

Gemma: Open Models Based on Gemini Research and Technology

Gemma Team, Thomas Mesnard, Cassidy Hardin +105

cs.LG2023

Measuring Forgetting of Memorized Training Examples

Matthew Jagielski, Om Thakkar, Florian Tramèr +8

cs.CR2023

Students Parrot Their Teachers: Membership Inference on Model Distillation

Matthew Jagielski, Milad Nasr, Christopher Choquette-Choo +2

cs.CY2023

Report of the 1st Workshop on Generative AI and Law

A. Feder Cooper, Katherine Lee, James Grimmelmann +32

cs.CL2023

MADLAD-400: A Multilingual And Document-Level Large Audited Dataset

Sneha Kudugunta, Isaac Caswell, Biao Zhang +8

stat.AP2018

Conditional regression based on a multivariate zero-inflated logistic normal model for microbiome relative abundance data

Zhigang Li, Katherine Lee, Margaret R. Karagas +4

cs.CL2023

PaLM 2 Technical Report

Rohan Anil, Andrew M. Dai, Orhan Firat +125

cs.CL2024

Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Gemini Team, Petko Georgiev, Ving Ian Lei +1132

cs.CL2022

PaLM: Scaling Language Modeling with Pathways

Aakanksha Chowdhery, Sharan Narang, Jacob Devlin +64

astro-ph.IM2024

The PLATO Mission

Heike Rauer, Conny Aerts, Juan Cabrera +842

stat.ML2022

What Does it Mean for a Language Model to Preserve Privacy?

Hannah Brown, Katherine Lee, Fatemehsadat Mireshghallah +2

cs.CL2025

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Gheorghe Comanici, Eric Bieber, Mike Schaekermann +3431

cs.CR2024

Stealing Part of a Production Language Model

Nicholas Carlini, Daniel Paleka, Krishnamurthy Dj Dvijotham +11

cs.LG2023

Preventing Verbatim Memorization in Language Models Gives a False Sense of Privacy

Daphne Ippolito, Florian Tramèr, Milad Nasr +5

cs.CY2024

Talkin' 'Bout AI Generation: Copyright and the Generative-AI Supply Chain

Katherine Lee, A. Feder Cooper, James Grimmelmann

cs.CL2023

A Pretrainer's Guide to Training Data: Measuring the Effects of Data Age, Domain Coverage, Quality, & Toxicity

Shayne Longpre, Gregory Yauney, Emily Reif +8

cs.LG2025

Measuring memorization in language models via probabilistic extraction

Jamie Hayes, Marika Swanberg, Harsh Chaudhari +6

cs.LG2023

Quantifying Memorization Across Neural Language Models

Nicholas Carlini, Daphne Ippolito, Matthew Jagielski +3

cs.CR2026

Exploring the limits of strong membership inference attacks on large language models

Jamie Hayes, Ilia Shumailov, Christopher A. Choquette-Choo +13

cs.CL2020

WT5?! Training Text-to-Text Models to Explain their Predictions

Sharan Narang, Colin Raffel, Katherine Lee +3

cs.LG2023

Scalable Extraction of Training Data from (Production) Language Models

Milad Nasr, Nicholas Carlini, Jonathan Hayase +7

cs.GT2024

An Abundance of Katherines: The Game Theory of Baby Naming

Katy Blumer, Kate Donahue, Katie Fritz +5

cs.CL2025

Privacy Ripple Effects from Adding or Removing Personal Information in Language Model Training

Jaydeep Borkar, Matthew Jagielski, Katherine Lee +3

astro-ph.SR2026

A method to derive self-consistent NLTE astrophysical parameters for 4 million high-resolution 4MOST stellar spectra in half a day with invertible neural networks

Victor F. Ksoll, Nicholas Storm, Maria Bergemann +5

cs.CL2025

Recite, Reconstruct, Recollect: Memorization in LMs as a Multifaceted Phenomenon

USVSN Sai Prashanth, Alvin Deng, Kyle O'Brien +9

cs.CL2024

Are aligned neural networks adversarially aligned?

Nicholas Carlini, Milad Nasr, Christopher A. Choquette-Choo +8

cs.CL2023

Counterfactual Memorization in Neural Language Models

Chiyuan Zhang, Daphne Ippolito, Katherine Lee +3

cs.LG2024

LMD3: Language Model Data Density Dependence

John Kirchenbauer, Garrett Honke, Gowthami Somepalli +5

cs.LG2024

Arbitrariness and Social Prediction: The Confounding Role of Variance in Fair Classification

A. Feder Cooper, Katherine Lee, Madiha Zahrah Choksi +6

cs.CL2026

Estimating near-verbatim extraction risk in language models with decoding-constrained beam search

A. Feder Cooper, Mark A. Lemley, Christopher De Sa +6