collaborators

7 papers

cs.CL2026

Are LLMs becoming similarly creative? Evidence from three years of models

Nirav Patel, Josiah Crossman, Eva Aggarwal +1

Many benchmarks track Large Language Model (LLM) performance on tasks with verifiable answers, but less is known about how LLM performance is evolving on open-ended tasks, where cr…

cs.CL2026

Where Models Converge and Humans Diverge: A Coverage Framework for Distributional Pluralism in Open-Ended Generation

Zini Yang, Emily Wenger, Richard So

When a large language model (LLM) writes Harry Potter fanfiction, it reliably produces fundamental elements of the Hogwarts universe, such as recognizable places and characters. Hu…

cs.CR2026

Identifying AI Web Scrapers Using Canary Tokens

Steven Seiden, Triss Ren, Caroline Zhang +3

From pre-training to query-time augmentation, web-scraped data helps to improve the quality and contextual relevancy of content generated by large language models (LLMs). However,…

cs.NI2025

Scrapers selectively respect robots.txt directives: evidence from a large-scale empirical study

Taein Kim, Karstan Bock, Claire Luo +3

Online data scraping has taken on new dimensions in recent years, as traditional scrapers have been joined by new AI-specific bots. To counteract unwanted scraping, many sites use…

cs.LG2025

What happens when generative AI models train recursively on each others' outputs?

Hung Anh Vu, Galen Reeves, Emily Wenger

The internet serves as a common source of training data for generative AI (genAI) models but is increasingly populated with AI-generated content. This duality raises the possibilit…

cs.LG2025

Causes and Consequences of Representational Similarity in Machine Learning Models

Zeyu Michael Li, Hung Anh Vu, Damilola Awofisayo +1

Numerous works have noted similarities in how machine learning models represent the world, even across modalities. Although much effort has been devoted to uncovering properties an…