3 papers
cs.CL2026
Beyond Public Access in LLM Pre-Training Data
Sruly Rosenblat, Tim O'Reilly, Ilan Strauss
Using a legally obtained dataset of 34 copyrighted O'Reilly Media books, we apply the DE-COP membership inference attack method to investigate whether OpenAI's large language model…
cs.DL2025
The Attribution Crisis in LLM Search Results
Ilan Strauss, Jangho Yang, Tim O'Reilly +2
Web-enabled LLMs frequently answer queries without crediting the web pages they consume, creating an "attribution gap" - the difference between relevant URLs read and those actuall…
cs.AI2025
Real-World Gaps in AI Governance Research
Ilan Strauss, Isobel Moure, Tim O'Reilly +1
Drawing on 1,178 safety and reliability papers from 9,439 generative AI papers (January 2020 - March 2025), we compare research outputs of leading AI companies (Anthropic, Google D…