1 paper · 1 filter
Sherin Muckatira, Jesse Geneson, Slava Gerovitch +3
Large language models have made substantial progress on mathematical reasoning, but existing benchmarks typically evaluate well-specified problems with final answers, step-by-step…