1 paper · 1 filter
Eldar Kurtic, Amir Moeini, Dan Alistarh
We introduce Mathador-LM, a new benchmark for evaluating the mathematical reasoning on large language models (LLMs), combining ruleset interpretation, planning, and problem-solving…