Six Fallacies in Substituting Large Language Models for Human Participants
arXiv:2402.04470 · doi:10.1177/25152459251357566
Abstract
Can AI systems like large language models (LLMs) replace human participants in behavioral and psychological research? Here I critically evaluate the "replacement" perspective and identify six interpretive fallacies that undermine its validity. These fallacies are: (1) equating token prediction with human intelligence, (2) treating LLMs as the average human, (3) interpreting alignment as explanation, (4) anthropomorphizing AI systems, (5) essentializing identities, and (6) substituting model data for human evidence. Each fallacy represents a potential misunderstanding about what LLMs are and what they can tell us about human cognition. The analysis distinguishes levels of similarity between LLMs and humans, particularly functional equivalence (outputs) versus mechanistic equivalence (processes), while highlighting both technical limitations (addressable through engineering) and conceptual limitations (arising from fundamental differences between statistical and biological intelligence). For each fallacy, specific safeguards are provided to guide responsible research practices. Ultimately, the analysis supports conceptualizing LLMs as pragmatic simulation tools--useful for role-play, rapid hypothesis testing, and computational modeling (provided their outputs are validated against human data)--rather than as replacements for human participants. This framework enables researchers to leverage language models productively while respecting the fundamental differences between machine intelligence and human thought.
References in corpus (17)
- Using cognitive psychology to understand GPT-3
- The Debate Over Understanding in AI's Large Language Models
- Cultural Bias and Cultural Alignment of Large Language Models
- Thinking Fast and Slow in Large Language Models
- Human-Like Intuitive Behavior and Reasoning Biases Emerged in Language Models -- and Disappeared in GPT-4
- From task structures to world models: What do LLMs know?
- Techniques for supercharging academic writing with generative AI
- The illusion of artificial inclusion
- Large Language Models Portray Socially Subordinate Groups as More Homogeneous, Consistent with a Bias Observed in Humans
- Language models align with human judgments on key grammatical constructions
- The Challenge of Using LLMs to Simulate Human Behavior: A Causal Inference Perspective
- Beyond principlism: Practical strategies for ethical AI use in research practices
- Evaluating Cognitive Maps and Planning in Large Language Models with CogEval
- Deanthropomorphising NLP: Can a Language Model Be Conscious?
- Large language models can replicate cross-cultural differences in personality
- Is Your LLM Outdated? A Deep Look at Temporal Generalization
- One fish, two fish, but not the whole sea: Alignment reduces language models' conceptual diversity