2 papers
cs.AI2026
Benchmark Everything Everywhere All at Once
Shiyun Xiong, Dongming Wu, Peiwen Sun +5
Benchmarks are fundamental for evaluating and advancing LLMs and MLLMs by providing standardized and explicit measures of performance. However, their construction is labor-intensiv…
cs.CL2025
UTMath: Math Evaluation with Unit Test via Reasoning-to-Coding Thoughts
Bo Yang, Qingping Yang, Yingwei Ma +1
The evaluation of mathematical reasoning capabilities is essential for advancing Artificial General Intelligence (AGI). While Large Language Models (LLMs) have shown impressive per…