Human heuristics for AI-generated language are flawed
arXiv:2206.07271 · doi:10.1073/pnas.2208839120
Abstract
Human communication is increasingly intermixed with language generated by AI. Across chat, email, and social media, AI systems suggest words, complete sentences, or produce entire conversations. AI-generated language is often not identified as such but presented as language written by humans, raising concerns about novel forms of deception and manipulation. Here, we study how humans discern whether verbal self-presentations, one of the most personal and consequential forms of language, were generated by AI. In six experiments, participants (N = 4,600) were unable to detect self-presentations generated by state-of-the-art AI language models in professional, hospitality, and dating contexts. A computational analysis of language features shows that human judgments of AI-generated language are hindered by intuitive but flawed heuristics such as associating first-person pronouns, use of contractions, or family topics with human-written language. We experimentally demonstrate that these heuristics make human judgment of AI-generated language predictable and manipulable, allowing AI systems to produce text perceived as "more human than human." We discuss solutions, such as AI accents, to reduce the deceptive potential of language generated by AI, limiting the subversion of human intuition.
References in corpus (4)
Cited by in corpus (19)
- Generative AI
- Co-Writing with Opinionated Language Models Affects Users' Views
- Can large language models replace humans in the systematic review process? Evaluating GPT-4's efficacy in screening and extracting data from peer-reviewed and grey literature in multiple languages
- A Design Space for Intelligent and Interactive Writing Assistants
- On the Conversational Persuasiveness of Large Language Models: A Randomized Controlled Trial
- Anatomy of an AI-powered malicious social botnet
- Comparing Sentence-Level Suggestions to Message-Level Suggestions in AI-Mediated Communication
- Evaluation of GPT-4 for chest X-ray impression generation: A reader study on performance and perception
- Help Me Reflect: Leveraging Self-Reflection Interface Nudges to Enhance Deliberativeness on Online Deliberation Platforms
- Position is Power: System Prompts as a Mechanism of Bias in Large Language Models (LLMs)
- Mapping the individual, social, and biospheric impacts of Foundation Models
- AI Rules? Characterizing Reddit Community Policies Towards AI-Generated Content
- Public Opinions About Copyright for AI-Generated Art: The Role of Egocentricity, Competition, and Experience
- Cognitive phantoms in LLMs through the lens of latent variables
- Should AI Mimic People? Understanding AI-Supported Writing Technology Among Black Users
- "There Has To Be a Lot That We're Missing": Moderating AI-Generated Content on Reddit
- Evaluation of Reliability Criteria for News Publishers with Large Language Models
- Humans can learn to detect AI-generated texts, or at least learn when they can't
- Detection Avoidance Techniques for Large Language Models