2 papers
cs.CL2025
Anticipatory Evaluation of Language Models
Jungsoo Park, Ethan Mendes, Gabriel Stanovsky +1
Progress in large language models is increasingly constrained by an evaluation bottleneck: benchmarks must be built and models run before iteration can begin. We investigate whethe…
cs.LG2025
Language Models can Self-Improve at State-Value Estimation for Better Search
Ethan Mendes, Alan Ritter
Collecting ground-truth rewards or human demonstrations for multi-step reasoning tasks is often prohibitively expensive, particularly in interactive domains such as web tasks. We i…