Out of One, Many: Using Language Models to Simulate Human Samples
arXiv:2209.06899 · doi:10.1017/pan.2023.2
Abstract
We propose and explore the possibility that language models can be studied as effective proxies for specific human sub-populations in social science research. Practical and research applications of artificial intelligence tools have sometimes been limited by problematic biases (such as racism or sexism), which are often treated as uniform properties of the models. We show that the "algorithmic bias" within one such tool -- the GPT-3 language model -- is instead both fine-grained and demographically correlated, meaning that proper conditioning will cause it to accurately emulate response distributions from a wide variety of human subgroups. We term this property "algorithmic fidelity" and explore its extent in GPT-3. We create "silicon samples" by conditioning the model on thousands of socio-demographic backstories from real human participants in multiple large surveys conducted in the United States. We then compare the silicon and human samples to demonstrate that the information contained in GPT-3 goes far beyond surface similarity. It is nuanced, multifaceted, and reflects the complex interplay between ideas, attitudes, and socio-cultural context that characterize human attitudes. We suggest that language models with sufficient algorithmic fidelity thus constitute a novel and powerful tool to advance understanding of humans and society across a variety of disciplines.
References in corpus (1)
Cited by in corpus (43)
- A Survey on Large Language Model based Autonomous Agents
- ChatGPT Outperforms Crowd-Workers for Text-Annotation Tasks
- Machine Culture
- EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined Criteria
- Generative Agent-Based Modeling: Unveiling Social System Dynamics through Coupling Mechanistic Models with Generative Artificial Intelligence
- The illusion of artificial inclusion
- The Challenge of Using LLMs to Simulate Human Behavior: A Causal Inference Perspective
- Detecting The Corruption Of Online Questionnaires By Artificial Intelligence
- Value Kaleidoscope: Engaging AI with Pluralistic Human Values, Rights, and Duties
- Can Large Language Models Capture Public Opinion about Global Warming? An Empirical Assessment of Algorithmic Fidelity and Bias
- "It's like a rubber duck that talks back": Understanding Generative AI-Assisted Data Analysis Workflows through a Participatory Prompting Study
- Strategic Behavior of Large Language Models: Game Structure vs. Contextual Framing
- Intelligent Computing Social Modeling and Methodological Innovations in Political Science in the Era of Large Language Models
- Human-centered NLP Fact-checking: Co-Designing with Fact-checkers using Matchmaking for AI
- Large-Language-Model-Powered Agent-Based Framework for Misinformation and Disinformation Research: Opportunities and Open Challenges
- Accuracy and Political Bias of News Source Credibility Ratings by Large Language Models
- MetaAgents: Large Language Model Based Agents for Decision-Making on Teaming
- On LLM Wizards: Identifying Large Language Models' Behaviors for Wizard of Oz Experiments
- Synthetically generated text for supervised text analysis
- Large Language Models' Accuracy in Emulating Human Experts' Evaluation of Public Sentiments about Heated Tobacco Products on Social Media
- Plurals: A System for Guiding LLMs Via Simulated Social Ensembles
- Judgment of Learning: A Human Ability Beyond Generative Artificial Intelligence
- Mapping and Influencing the Political Ideology of Large Language Models using Synthetic Personas
- Helpful assistant or fruitful facilitator? Investigating how personas affect language model behavior
- Vox Populi, Vox AI? Using Language Models to Estimate German Public Opinion
- LLM-Agent-UMF: LLM-based Agent Unified Modeling Framework for Seamless Design of Multi Active/Passive Core-Agent Architectures
- Bootstrapping LLM-based Task-Oriented Dialogue Agents via Self-Talk
- Exploring LLMs for Automated Generation and Adaptation of Questionnaires
- Adaptive political surveys and GPT-4: Tackling the cold start problem with simulated user interactions
- Creating and Evaluating Personas Using Generative AI: A Scoping Review of 81 Articles
- Financial Stability Implications of Generative AI: Taming the Animal Spirits
- Redefining Research Crowdsourcing: Incorporating Human Feedback with LLM-Powered Digital Twins
- Multimodal Fusion with LLMs for Engagement Prediction in Natural Conversation
- Opinion dynamics model of collaborative learning
- Valuing Time in Silicon: Can Large Language Models Replicate Human Value of Travel Time
- DxHF: Providing High-Quality Human Feedback for LLM Alignment via Interactive Decomposition
- Shiny Stories, Hidden Struggles: Investigating the Representation of Disability Through the Lens of LLMs
- AI agents can coordinate beyond human scale
- Embracing Dialectic Intersubjectivity: Coordination of Different Perspectives in Content Analysis with LLM Persona Simulation
- "Amazing, They All Lean Left" -- Analyzing the Political Temperaments of Current LLMs
- What does AI consider praiseworthy?
- Political Ideology Shifts in Large Language Models
- Evaluating the Effectiveness of Persona Simulation in Opinion Prediction with GPT-4.1