Large Language Models are Zero-Shot Reasoners
arXiv:2205.11916
Abstract
Pretrained large language models (LLMs) are widely used in many sub-fields of natural language processing (NLP) and generally known as excellent few-shot learners with task-specific exemplars. Notably, chain of thought (CoT) prompting, a recent technique for eliciting complex multi-step reasoning through step-by-step answer examples, achieved the state-of-the-art performances in arithmetics and symbolic reasoning, difficult system-2 tasks that do not follow the standard scaling laws for LLMs. While these successes are often attributed to LLMs' ability for few-shot learning, we show that LLMs are decent zero-shot reasoners by simply adding "Let's think step by step" before each answer. Experimental results demonstrate that our Zero-shot-CoT, using the same single prompt template, significantly outperforms zero-shot LLM performances on diverse benchmark reasoning tasks including arithmetics (MultiArith, GSM8K, AQUA-RAT, SVAMP), symbolic reasoning (Last Letter, Coin Flip), and other logical reasoning tasks (Date Understanding, Tracking Shuffled Objects), without any hand-crafted few-shot examples, e.g. increasing the accuracy on MultiArith from 17.7% to 78.7% and GSM8K from 10.4% to 40.7% with large InstructGPT model (text-davinci-002), as well as similar magnitudes of improvements with another off-the-shelf large model, 540B parameter PaLM. The versatility of this single prompt across very diverse reasoning tasks hints at untapped and understudied fundamental zero-shot capabilities of LLMs, suggesting high-level, multi-task broad cognitive capabilities may be extracted by simple prompting. We hope our work not only serves as the minimal strongest zero-shot baseline for the challenging reasoning benchmarks, but also highlights the importance of carefully exploring and analyzing the enormous zero-shot knowledge hidden inside LLMs before crafting finetuning datasets or few-shot exemplars.
Accepted to NeurIPS2022. Our code is available at https://github.com/kojima-takeshi188/zero_shot_cot
Cited by in corpus (72)
- Automatic Generation of Programming Exercises and Code Explanations using Large Language Models
- Unleashing the potential of prompt engineering for large language models
- Large AI Models in Health Informatics: Applications, Challenges, and the Future
- Auditing large language models: a three-layered approach
- Bad Actor, Good Advisor: Exploring the Role of Large Language Models in Fake News Detection
- Uncovering ChatGPT's Capabilities in Recommender Systems
- Towards Interpretable Mental Health Analysis with Large Language Models
- How understanding large language models can inform the use of ChatGPT in physics education
- A Quantitative and Qualitative Evaluation of LLM-Based Explainable Fault Localization
- Generative Artificial Intelligence in Small and Medium Enterprises: Navigating its Promises and Challenges
- Towards autonomous system: flexible modular production system enhanced with large language model agents
- DALL-E-Bot: Introducing Web-Scale Diffusion Models to Robotics
- What Makes Good In-context Demonstrations for Code Intelligence Tasks with LLMs?
- Deception Abilities Emerged in Large Language Models
- Large Language Models as Zero-Shot Human Models for Human-Robot Interaction
- UAVs Meet LLMs: Overviews and Perspectives Toward Agentic Low-Altitude Mobility
- Using LLMs in Software Requirements Specifications: An Empirical Evaluation
- The Impact of AI in Physics Education: A Comprehensive Review from GCSE to University Levels
- Real-World Robot Applications of Foundation Models: A Review
- A Practical Survey on Zero-shot Prompt Design for In-context Learning
- Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models
- AppPoet: Large Language Model based Android malware detection via multi-view prompt engineering
- LINC: A Neurosymbolic Approach for Logical Reasoning by Combining Language Models with First-Order Logic Provers
- Generative Speech Recognition Error Correction with Large Language Models and Task-Activating Prompting
- Evaluation of LLM Chatbots for OSINT-based Cyber Threat Awareness
- ThoughtSource: A central hub for large language model reasoning data
- Exploring the Roles of Large Language Models in Reshaping Transportation Systems: A Survey, Framework, and Roadmap
- OntoChatGPT Information System: Ontology-Driven Structured Prompts for ChatGPT Meta-Learning
- Could ChatGPT get an Engineering Degree? Evaluating Higher Education Vulnerability to AI Assistants
- Developing ChemDFM as a large language foundation model for chemistry
- Examining Inter-Consistency of Large Language Models Collaboration: An In-depth Analysis via Debate
- Foundation Models for Geospatial Reasoning: Assessing Capabilities of Large Language Models in Understanding Geometries and Topological Spatial Relations
- Prompted LLMs as Chatbot Modules for Long Open-domain Conversation
- Domain-specific ChatBots for Science using Embeddings
- Embedding Democratic Values into Social Media AIs via Societal Objective Functions
- Improving deep learning with prior knowledge and cognitive models: A survey on enhancing explainability, adversarial robustness and zero-shot learning
- Human I/O: Towards a Unified Approach to Detecting Situational Impairments
- Evaluation of GPT and BERT-based models on identifying protein-protein interactions in biomedical text
- Joint Estimation and Prediction of City-wide Delivery Demand: A Large Language Model Empowered Graph-based Learning Approach
- Theory of Mind abilities of Large Language Models in Human-Robot Interaction : An Illusion?
- Solving Math Word Problems via Cooperative Reasoning induced Language Models
- Out of the Cage: How Stochastic Parrots Win in Cyber Security Environments
- Large Language Models for Scholarly Ontology Generation: An Extensive Analysis in the Engineering Field
- LLM-Mediated Domain-Specific Voice Agents: The Case of TextileBot
- A Survey on Knowledge Organization Systems of Research Fields: Resources and Challenges
- On Sarcasm Detection with OpenAI GPT-based Models
- Eight challenges in developing theory of intelligence
- ScoNe: Benchmarking Negation Reasoning in Language Models With Fine-Tuning and In-Context Learning
- Cases of EFL Secondary Students' Prompt Engineering Pathways to Complete a Writing Task with ChatGPT
- HiPrompt: Few-Shot Biomedical Knowledge Fusion via Hierarchy-Oriented Prompting
- An exploration of features to improve the generalisability of fake news detection models
- Large Language Models for Combinatorial Optimization: A Systematic Review
- Translating Legalese: Enhancing Public Understanding of Court Opinions with Legal Summarizers
- The Problem of Atypicality in LLM-Powered Psychiatry
- Beyond Self-Consistency: Ensemble Reasoning Boosts Consistency and Accuracy of LLMs in Cancer Staging
- "Oh LLM, I'm Asking Thee, Please Give Me a Decision Tree": Zero-Shot Decision Tree Induction and Embedding with Large Language Models
- Complex QA and language models hybrid architectures, Survey
- Relational Programming with Foundation Models
- Piloting Copilot, Codex, and StarCoder2: Hot Temperature, Cold Prompts, or Black Magic?
- Prompt engineering for bibliographic web-scraping
- OPT-R: Exploring the Role of Explanations in Finetuning and Prompting for Reasoning Skills of Large Language Models
- Concept-Guided Chain-of-Thought Prompting for Pairwise Comparison Scoring of Texts with Large Language Models
- Large Language Models and Multimodal Retrieval for Visual Word Sense Disambiguation
- Evaluating Large Language Models in Code Generation: INFINITE Methodology for Defining the Inference Index
- Analyzing FOMC Minutes: Accuracy and Constraints of Language Models
- Flows: Building Blocks of Reasoning and Collaborating AI
- Deep Natural Language Feature Learning for Interpretable Prediction
- Automated Theorem Provers Help Improve Large Language Model Reasoning
- Do GPT Language Models Suffer From Split Personality Disorder? The Advent Of Substrate-Free Psychometrics
- Large Language Models are biased to overestimate profoundness
- A Hierarchical Framework for Measuring Scientific Paper Innovation via Large Language Models
- Cheap Learning: Maximising Performance of Language Models for Social Data Science Using Minimal Data