8 citations · 8 across the 6 of their papers we have counts for
1 paper · 1 filter
Isabelle Lee, Sarah Liaw, Dani Yogatama
Reasoning in language models is difficult to evaluate: natural-language traces are unverifiable, symbolic datasets are too small, and most benchmarks conflate heuristics with infer…