1 citations · 1 across the 3 of their papers we have counts for
4 papers · 1 filter
Generating Data-Driven Reasoning Rubrics for Domain-Adaptive Reward Modeling
Kate Sanders, Nathaniel Weir, Sapana Chaudhary +2
An impediment to using Large Language Models (LLMs) for reasoning output verification is that LLMs struggle to reliably identify errors in thinking traces, particularly in long out…
Natural Language Deduction with Incomplete Information
Zayne Sprague, Kaj Bostrom, Swarat Chaudhuri +1
A growing body of work studies how to answer a question or verify a claim by generating a natural language "proof": a chain of deductive inferences yielding the answer based on a s…
Flexible Generation of Natural Language Deductions
Kaj Bostrom, Xinyu Zhao, Swarat Chaudhuri +1
An interpretable system for open-domain reasoning needs to express its reasoning process in a transparent form. Natural language is an attractive representation for this purpose --…
Byte Pair Encoding is Suboptimal for Language Model Pretraining
Kaj Bostrom, Greg Durrett
The success of pretrained transformer language models (LMs) in natural language processing has led to a wide range of pretraining setups. In particular, these models employ a varie…