2 papers
cs.CL2026
Generating Data-Driven Reasoning Rubrics for Domain-Adaptive Reward Modeling
Kate Sanders, Nathaniel Weir, Sapana Chaudhary +2
An impediment to using Large Language Models (LLMs) for reasoning output verification is that LLMs struggle to reliably identify errors in thinking traces, particularly in long out…
cs.LG2025
Randomly Sampled Language Reasoning Problems Elucidate Limitations of In-Context Learning
Kavi Gupta, Kate Sanders, Armando Solar-Lezama
While LLMs have revolutionized the field of machine learning due to their high performance on a strikingly wide range of problems, they are also known to hallucinate false answers…