Deception Abilities Emerged in Large Language Models
arXiv:2307.16513 · doi:10.1073/pnas.2317967121
Abstract
Large language models (LLMs) are currently at the forefront of intertwining artificial intelligence (AI) systems with human communication and everyday life. Thus, aligning them with human values is of great importance. However, given the steady increase in reasoning abilities, future LLMs are under suspicion of becoming able to deceive human operators and utilizing this ability to bypass monitoring efforts. As a prerequisite to this, LLMs need to possess a conceptual understanding of deception strategies. This study reveals that such strategies emerged in state-of-the-art LLMs, such as GPT-4, but were non-existent in earlier LLMs. We conduct a series of experiments showing that state-of-the-art LLMs are able to understand and induce false beliefs in other agents, that their performance in complex deception scenarios can be amplified utilizing chain-of-thought reasoning, and that eliciting Machiavellianism in LLMs can alter their propensity to deceive. In sum, revealing hitherto unknown machine behavior in LLMs, our study contributes to the nascent field of machine psychology.
References in corpus (12)
- Sparks of Artificial General Intelligence: Early experiments with GPT-4
- Scaling Instruction-Finetuned Language Models
- Human-Like Intuitive Behavior and Reasoning Biases Emerged in Language Models -- and Disappeared in GPT-4
- How is ChatGPT's behavior changing over time?
- Deception Abilities Emerged in Large Language Models
- Jailbroken: How Does LLM Safety Training Fail?
- An Overview of Catastrophic AI Risks
- Alignment of Language Agents
- Boosting Theory-of-Mind Performance in Large Language Models via Prompting
- Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark
- DERA: Enhancing Large Language Model Completions with Dialog-Enabled Resolving Agents
- Deceptive Alignment Monitoring
Cited by in corpus (12)
- Mapping the Ethics of Generative AI: A Comprehensive Scoping Review
- Deception Abilities Emerged in Large Language Models
- Do LLMs write like humans? Variation in grammatical and rhetorical styles
- Exploring the Frontiers of LLMs in Psychological Applications: A Comprehensive Review
- Deception and Manipulation in Generative AI
- Theory of Mind abilities of Large Language Models in Human-Robot Interaction : An Illusion?
- VelLMes: A high-interaction AI-based deception framework
- Consumer Law for AI Agents
- Hallmarks of Deception in Asset-Exchange Models
- Neural Decompiling of Tracr Transformers
- When Do Large Language Models Exhibit Unsolicited Deception?
- "Amazing, They All Lean Left" -- Analyzing the Political Temperaments of Current LLMs