5 citations · 24 across the 17 of their papers we have counts for
14 papers · 1 filter
AI Fiction in the Wild
Neel Gupta, Maria Antoniak, Melanie Walsh
Some professional authors are beginning to use AI tools to help produce their fiction writing. Are readers using AI to generate fiction, too? Drawing on over 500,000 anonymized, En…
Characterizing Narrative Content in Web-scale LLM Pretraining Data
Teagan Johnson, Elliott Ash, Andrew Piper +1
The narrative composition of web-scale LLM pretraining corpora remains largely unexplored, even though narrative is a fundamental mode of human communication. We present the first…
What and Whose Knowledge? Measuring Epistemic Diversity in Large Language Models
Dustin Wright, Sarah Masud, Jared Moore +5
Large language models (LLMs) are increasingly used as primary knowledge sources, yet their epistemic diversity - defined as the diversity of real-world claims in their outputs - ha…
Research Borderlands: Analysing Writing Across Research Cultures
Shaily Bhatt, Tal August, Maria Antoniak
Improving cultural competence of language technologies is important. However most recent works rarely engage with the communities they study, and instead rely on synthetic setups a…
Meaning Is Not A Metric: Using LLMs to make cultural context legible at scale
Cody Kommers, Drew Hemment, Maria Antoniak +4
This position paper argues that large language models (LLMs) can make cultural context, and therefore human meaning, legible at an unprecedented scale in AI-based sociotechnical sy…
Information-Guided Identification of Training Data Imprint in (Proprietary) Large Language Models
Abhilasha Ravichander, Jillian Fisher, Taylor Sorensen +6
High-quality training data has proven crucial for developing performant large language models (LLMs). However, commercial LLM providers disclose few, if any, details about the data…