1 paper · 1 filter
Isabelle Lee, Sarah Liaw, Dani Yogatama
Reasoning in language models is difficult to evaluate: natural-language traces are unverifiable, symbolic datasets are too small, and most benchmarks conflate heuristics with infer…