Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
Scaling Test-time Compute for LLM Agents
King Zhu, Hanhao Li, Siwei Wu +12
Scaling test time compute has shown remarkable success in improving the reasoning abilities of large language models (LLMs). In this work, we conduct the first systematic explorati…
cs.AI2025
DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models
Xiaoyang Chen, Xinan Dai, Yu Du +28
To advance the mathematical proficiency of large language models (LLMs), the DeepMath team has launched an open-source initiative aimed at developing an open mathematical LLM and s…