1 paper · 1 filter
Jiayu Fu, Mourad Heddaya, Chenhao Tan
Numerous math benchmarks exist to evaluate LLMs' mathematical capabilities. However, most involve extensive manual effort and are difficult to scale. Consequently, they cannot keep…