12 papers
AI Fiction in the Wild
Neel Gupta, Maria Antoniak, Melanie Walsh
Some professional authors are beginning to use AI tools to help produce their fiction writing. Are readers using AI to generate fiction, too? Drawing on over 500,000 anonymized, En…
Characterizing Narrative Content in Web-scale LLM Pretraining Data
Teagan Johnson, Elliott Ash, Andrew Piper +1
The narrative composition of web-scale LLM pretraining corpora remains largely unexplored even though narrative is a fundamental mode of human communication. We present the first f…
Computational Hermeneutics: Evaluating generative AI as a cultural technology
Cody Kommers, Ruth Ahnert, Maria Antoniak +35
Generative AI systems are increasingly recognized as cultural technologies, yet current evaluation frameworks often treat culture as a variable to be measured rather than fundament…
Epistemic Diversity and Knowledge Collapse in Large Language Models
Dustin Wright, Sarah Masud, Jared Moore +5
Large language models (LLMs) tend to generate homogenous texts, which may impact the diversity of knowledge generated across different outputs. Given their potential to replace exi…
Meaning Is Not A Metric: Using LLMs to make cultural context legible at scale
Cody Kommers, Drew Hemment, Maria Antoniak +4
This position paper argues that large language models (LLMs) can make cultural context, and therefore human meaning, legible at an unprecedented scale in AI-based sociotechnical sy…
LLM-Generated or Human-Written? Comparing Review and Non-Review Papers on ArXiv
Yanai Elazar, Maria Antoniak
ArXiv recently prohibited the upload of unpublished review papers to its servers in the Computer Science domain, citing a high prevalence of LLM-generated content in these categori…