LINC: A Neurosymbolic Approach for Logical Reasoning by Combining Language Models with First-Order Logic Provers
arXiv:2310.15164 · doi:10.18653/v1/2023.emnlp-main.313
Abstract
Logical reasoning, i.e., deductively inferring the truth value of a conclusion from a set of premises, is an important task for artificial intelligence with wide potential impacts on science, mathematics, and society. While many prompting-based strategies have been proposed to enable Large Language Models (LLMs) to do such reasoning more effectively, they still appear unsatisfactory, often failing in subtle and unpredictable ways. In this work, we investigate the validity of instead reformulating such tasks as modular neurosymbolic programming, which we call LINC: Logical Inference via Neurosymbolic Computation. In LINC, the LLM acts as a semantic parser, translating premises and conclusions from natural language to expressions in first-order logic. These expressions are then offloaded to an external theorem prover, which symbolically performs deductive inference. Leveraging this approach, we observe significant performance gains on FOLIO and a balanced subset of ProofWriter for three different models in nearly all experimental conditions we evaluate. On ProofWriter, augmenting the comparatively small open-source StarCoder+ (15.5B parameters) with LINC even outperforms GPT-3.5 and GPT-4 with Chain-of-Thought (CoT) prompting by an absolute 38% and 10%, respectively. When used with GPT-4, LINC scores 26% higher than CoT on ProofWriter while performing comparatively on FOLIO. Further analysis reveals that although both methods on average succeed roughly equally often on this dataset, they exhibit distinct and complementary failure modes. We thus provide promising evidence for how logical reasoning over natural language can be tackled through jointly leveraging LLMs alongside symbolic provers. All corresponding code is publicly available at https://github.com/benlipkin/linc
EMNLP Main 2023 (Outstanding Paper Award)
References in corpus (33)
- Training language models to follow instructions with human feedback
- PaLM: Scaling Language Modeling with Pathways
- Large Language Models are Zero-Shot Reasoners
- LaMDA: Language Models for Dialog Applications
- Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
- Toolformer: Language Models Can Teach Themselves to Use Tools
- Self-Refine: Iterative Refinement with Self-Feedback
- StarCoder: may the source be with you!
- A Neural Network Solves, Explains, and Generates University Math Problems by Program Synthesis and Few-Shot Learning at Human Level
- The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only
- Check Your Facts and Try Again: Improving Large Language Models with External Knowledge and Automated Feedback
- Augmented Language Models: a Survey
- Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks
- Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning
- PAL: Program-aided Language Models
- Teaching Large Language Models to Self-Debug
- Faith and Fate: Limits of Transformers on Compositionality
- Synchromesh: Reliable code generation from pre-trained language models
- Efficient Training of Language Models to Fill in the Middle
- Exploring Length Generalization in Large Language Models
- The Stack: 3 TB of permissively licensed source code
- Compositional Semantic Parsing with Large Language Models
- Large Language Models Cannot Self-Correct Reasoning Yet
- Binding Language Models in Symbolic Languages
- FOLIO: Natural Language Reasoning with First-Order Logic
- Mind's Eye: Grounded Language Model Reasoning through Simulation
- Language Model Cascades
- Testing the General Deductive Reasoning Capacity of Large Language Models Using OOD Examples
- ProofNet: Autoformalizing and Formally Proving Undergraduate-Level Mathematics
- Formal Specifications from Natural Language
- A Generalist Neural Algorithmic Learner
- nl2spec: Interactively Translating Unstructured Natural Language to Temporal Logics with Large Language Models
- Is Self-Repair a Silver Bullet for Code Generation?