1 paper
Jicheng Ma, Guohua Wang, Xinhua Feng +3
Current evaluations of mathematical reasoning in large language models (LLMs) are dominated by static benchmarks, either derived from competition-style problems or curated through…