1 paper
Yan Liu, Renren Jin, Ling Shi +2
To thoroughly assess the mathematical reasoning abilities of Large Language Models (LLMs), we need to carefully curate evaluation datasets covering diverse mathematical concepts an…