Self-Consistency Improves Chain of Thought Reasoning in Language Models
arXiv:2203.11171
Abstract
Chain-of-thought prompting combined with pre-trained large language models has achieved encouraging results on complex reasoning tasks. In this paper, we propose a new decoding strategy, self-consistency, to replace the naive greedy decoding used in chain-of-thought prompting. It first samples a diverse set of reasoning paths instead of only taking the greedy one, and then selects the most consistent answer by marginalizing out the sampled reasoning paths. Self-consistency leverages the intuition that a complex reasoning problem typically admits multiple different ways of thinking leading to its unique correct answer. Our extensive empirical evaluation shows that self-consistency boosts the performance of chain-of-thought prompting with a striking margin on a range of popular arithmetic and commonsense reasoning benchmarks, including GSM8K (+17.9%), SVAMP (+11.0%), AQuA (+12.2%), StrategyQA (+6.4%) and ARC-challenge (+3.9%).
Published at ICLR 2023. V2: added PaLM results; V3: added UL2 results; V4: camera ready version at ICLR 2023
Cited by in corpus (45)
- Unifying Large Language Models and Knowledge Graphs: A Roadmap
- Graph of Thoughts: Solving Elaborate Problems with Large Language Models
- Recommender Systems in the Era of Large Language Models (LLMs)
- AI Agents Under Threat: A Survey of Key Security Challenges and Future Pathways
- OpenMedLM: Prompt engineering can out-perform fine-tuning in medical question-answering with open-source large language models
- How understanding large language models can inform the use of ChatGPT in physics education
- Large Language Model Is Not a Good Few-shot Information Extractor, but a Good Reranker for Hard Samples!
- A Quantitative and Qualitative Evaluation of LLM-Based Explainable Fault Localization
- LLM for SoC Security: A Paradigm Shift
- The Impact of AI in Physics Education: A Comprehensive Review from GCSE to University Levels
- A Large Language Model Approach to Educational Survey Feedback Analysis
- GPT-4 can pass the Korean National Licensing Examination for Korean Medicine Doctors
- AppPoet: Large Language Model based Android malware detection via multi-view prompt engineering
- Enhancing Knowledge Retrieval with In-Context Learning and Semantic Search through Generative AI
- CoAnnotating: Uncertainty-Guided Work Allocation between Human and Large Language Models for Data Annotation
- Examining Inter-Consistency of Large Language Models Collaboration: An In-depth Analysis via Debate
- RELIC: Investigating Large Language Model Responses using Self-Consistency
- Automated Review Generation Method Based on Large Language Models
- REFINER: Reasoning Feedback on Intermediate Representations
- Exploring the Sensitivity of LLMs' Decision-Making Capabilities: Insights from Prompt Variation and Hyperparameters
- T cell receptor binding prediction: A machine learning revolution
- Large Knowledge Model: Perspectives and Challenges
- Quantum Natural Language Processing
- Intelligent Computing Social Modeling and Methodological Innovations in Political Science in the Era of Large Language Models
- Theory of Mind abilities of Large Language Models in Human-Robot Interaction : An Illusion?
- Solving Math Word Problems via Cooperative Reasoning induced Language Models
- Large Language Models for Scholarly Ontology Generation: An Extensive Analysis in the Engineering Field
- Do PLMs Know and Understand Ontological Knowledge?
- GenQREnsemble: Zero-Shot LLM Ensemble Prompting for Generative Query Reformulation
- Guideline Learning for In-context Information Extraction
- LLM Agent Framework for Intelligent Change Analysis in Urban Environment using Remote Sensing Imagery
- Exploring the Landscape of Natural Language Processing Research
- Cross-Lingual Consistency of Factual Knowledge in Multilingual Language Models
- A Meta-Evaluation of Faithfulness Metrics for Long-Form Hospital-Course Summarization
- Diverse Title Generation for Stack Overflow Posts with Multiple Sampling Enhanced Transformer
- Beyond Self-Consistency: Ensemble Reasoning Boosts Consistency and Accuracy of LLMs in Cancer Staging
- Relational Programming with Foundation Models
- Piloting Copilot, Codex, and StarCoder2: Hot Temperature, Cold Prompts, or Black Magic?
- Complex QA and language models hybrid architectures, Survey
- Deep Natural Language Feature Learning for Interpretable Prediction
- Explainability Through Systematicity: The Hard Systematicity Challenge for Artificial Intelligence
- CoRRPUS: Code-based Structured Prompting for Neurosymbolic Story Understanding
- HELIOS: Harmonizing Early Fusion, Late Fusion, and LLM Reasoning for Multi-Granular Table-Text Retrieval
- End-to-End Bangla AI for Solving Math Olympiad Problem Benchmark: Leveraging Large Language Model Using Integrated Approach
- Simple and Effective Input Reformulations for Translation