7 papers
Are LLMs becoming similarly creative? Evidence from three years of models
Nirav Patel, Josiah Crossman, Eva Aggarwal +1
Many benchmarks track Large Language Model (LLM) performance on tasks with verifiable answers, but less is known about how LLM performance is evolving on open-ended tasks, where cr…
Where Models Converge and Humans Diverge: A Coverage Framework for Distributional Pluralism in Open-Ended Generation
Zini Yang, Emily Wenger, Richard So
When a large language model (LLM) writes Harry Potter fanfiction, it reliably produces fundamental elements of the Hogwarts universe, such as recognizable places and characters. Hu…
Identifying AI Web Scrapers Using Canary Tokens
Steven Seiden, Triss Ren, Caroline Zhang +3
From pre-training to query-time augmentation, web-scraped data helps to improve the quality and contextual relevancy of content generated by large language models (LLMs). However,…
Scrapers selectively respect robots.txt directives: evidence from a large-scale empirical study
Taein Kim, Karstan Bock, Claire Luo +3
Online data scraping has taken on new dimensions in recent years, as traditional scrapers have been joined by new AI-specific bots. To counteract unwanted scraping, many sites use…
What happens when generative AI models train recursively on each others' outputs?
Hung Anh Vu, Galen Reeves, Emily Wenger
The internet serves as a common source of training data for generative AI (genAI) models but is increasingly populated with AI-generated content. This duality raises the possibilit…
Causes and Consequences of Representational Similarity in Machine Learning Models
Zeyu Michael Li, Hung Anh Vu, Damilola Awofisayo +1
Numerous works have noted similarities in how machine learning models represent the world, even across modalities. Although much effort has been devoted to uncovering properties an…