Exploring the psychology of LLMs' Moral and Legal Reasoning
arXiv:2308.01264 · doi:10.1016/j.artint.2024.104145
Abstract
Large language models (LLMs) exhibit expert-level performance in tasks across a wide range of different domains. Ethical issues raised by LLMs and the need to align future versions makes it important to know how state of the art models reason about moral and legal issues. In this paper, we employ the methods of experimental psychology to probe into this question. We replicate eight studies from the experimental literature with instances of Google's Gemini Pro, Anthropic's Claude 2.1, OpenAI's GPT-4, and Meta's Llama 2 Chat 70b. We find that alignment with human responses shifts from one experiment to another, and that models differ amongst themselves as to their overall alignment, with GPT-4 taking a clear lead over all other models we tested. Nonetheless, even when LLM-generated responses are highly correlated to human responses, there are still systematic differences, with a tendency for models to exaggerate effects that are present among humans, in part by reducing variance. This recommends caution with regards to proposals of replacing human participants with current state-of-the-art LLMs in psychological research and highlights the need for further research about the distinctive aspects of machine psychology.
References in corpus (10)
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Sparks of Artificial General Intelligence: Early experiments with GPT-4
- Gemini: A Family of Highly Capable Multimodal Models
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Constitutional AI: Harmlessness from AI Feedback
- Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models
- Whose Opinions Do Language Models Reflect?
- Machine Psychology
- MoCa: Measuring Human-Language Model Alignment on Causal and Moral Judgment Tasks
- Specific versus General Principles for Constitutional AI
Cited by in corpus (4)
- Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence
- Exploring the Frontiers of LLMs in Psychological Applications: A Comprehensive Review
- Metaheuristics and Large Language Models Join Forces: Toward an Integrated Optimization Approach
- The Epistemological Consequences of Large Language Models: Rethinking collective intelligence and institutional knowledge