From Prompts to Constructs: A Dual-Validity Framework for Large Language Model Research in Psychology
arXiv:2506.16697 · doi:10.1146/annurev-psych-100925-034807
Abstract
Large language models (LLMs) are entering psychological research both as tools and as objects of inquiry. Yet many studies apply human instruments to LLMs without establishing that the outputs are reliable or interpretable, raising the risk of measurement phantoms--statistical regularities mistaken for genuine psychological phenomena. This review argues that robust AI psychological research requires integrating two methodological traditions: psychometric validation of what a score means and causal inference standards for what the results warrant. It develops a dual-validity framework in which evidentiary demands scale with scientific ambition: from tool use through behavioral characterization and human simulation to cognitive modeling. Classifying text may require only accuracy and reliability; claiming that an LLM simulates anxiety or illuminates cognitive mechanisms requires additional evidence, including construct validity evidence and experimental controls. Progress depends on developing computational analogs of psychological constructs rather than assuming human measures automatically apply to language models.
References in corpus (31)
- Using cognitive psychology to understand GPT-3
- Evaluating Large Language Models in Theory of Mind Tasks
- Towards Understanding Sycophancy in Language Models
- How to avoid machine learning pitfalls: a guide for academic researchers
- Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies
- Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
- Techniques for supercharging academic writing with generative AI
- Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
- The Moral Machine Experiment on Large Language Models
- Can Large Language Models Capture Public Opinion about Global Warming? An Empirical Assessment of Algorithmic Fidelity and Bias
- Does Prompt Formatting Have Any Impact on LLM Performance?
- Who is GPT-3? An Exploration of Personality, Values and Demographics
- A Philosophical Introduction to Language Models - Part II: The Way Forward
- Large Language Models and Cognitive Science: A Comprehensive Review of Similarities, Differences, and Challenges
- Large-scale moral machine experiment on large language models
- Large Language Model Psychometrics: A Systematic Review of Evaluation, Validation, and Enhancement
- Evaluating Large Language Models with Psychometrics
- Cognitive phantoms in LLMs through the lens of latent variables
- Take Caution in Using LLMs as Human Surrogates: Scylla Ex Machina
- Challenging the Validity of Personality Tests for Large Language Models
- AI Psychometrics: Evaluating the Psychological Reasoning of Large Language Models with Psychometric Validities
- Self-Assessment Tests are Unreliable Measures of LLM Personality
- Large Language Models Show Human-like Social Desirability Biases in Survey Responses
- Do LLMs Have Distinct and Consistent Personality? TRAIT: Personality Testset designed for LLMs with Psychometrics
- Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots
- A validity-guided workflow for robust large language model research in psychology
- GPT-ology, Computational Models, Silicon Sampling: How should we think about LLMs in Cognitive Science?
- Language Models Are Capable of Metacognitive Monitoring and Control of Their Internal Activations
- Beyond BFI: The CSI for Enhanced Reliability and Validity in Evaluating LLM Personality Traits
- Assessing Social Alignment: Do Personality-Prompted Large Language Models Behave Like Humans?
- Sense and Sensitivity: Evaluating the simulation of social dynamics via Large Language Models