REFINER: Reasoning Feedback on Intermediate Representations
arXiv:2304.01904
Abstract
Language models (LMs) have recently shown remarkable performance on reasoning tasks by explicitly generating intermediate inferences, e.g., chain-of-thought prompting. However, these intermediate inference steps may be inappropriate deductions from the initial context and lead to incorrect final predictions. Here we introduce REFINER, a framework for finetuning LMs to explicitly generate intermediate reasoning steps while interacting with a critic model that provides automated feedback on the reasoning. Specifically, the critic provides structured feedback that the reasoning LM uses to iteratively improve its intermediate arguments. Empirical evaluations of REFINER on three diverse reasoning tasks show significant improvements over baseline LMs of comparable scale. Furthermore, when using GPT-3.5 or ChatGPT as the reasoner, the trained critic significantly improves reasoning without finetuning the reasoner. Finally, our critic model is trained without expensive human-in-the-loop data but can be substituted with humans at inference time.
Accepted at EACL 2024
References in corpus (15)
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- ReAct: Synergizing Reasoning and Acting in Language Models
- Constitutional AI: Harmlessness from AI Feedback
- Reflexion: Language Agents with Verbal Reinforcement Learning
- Self-Refine: Iterative Refinement with Self-Feedback
- Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy Optimization
- The Unreliability of Explanations in Few-shot Prompting for Textual Reasoning
- Self-critiquing models for assisting human evaluators
- Generating Sequences by Learning to Self-Correct
- ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning
- Large Language Models Can Self-Improve
- Training Language Models with Language Feedback
- Tell me why! Explanations support learning relational and causal structure
- Argumentative Reward Learning: Reasoning About Human Preferences
Cited by in corpus (4)
- Graph of Thoughts: Solving Elaborate Problems with Large Language Models
- Could ChatGPT get an Engineering Degree? Evaluating Higher Education Vulnerability to AI Assistants
- Can ChatGPT Perform Reasoning Using the IRAC Method in Analyzing Legal Scenarios Like a Lawyer?
- Flows: Building Blocks of Reasoning and Collaborating AI