1 paper
Xiaoxuan Wang, Ziniu Hu, Pan Lu +7
Most of the existing Large Language Model (LLM) benchmarks on scientific problem reasoning focus on problems grounded in high-school subjects and are confined to elementary algebra…