Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
PyraMathBench: Evaluating and Improving Mathematical Capability in Large Language Models
Zetian Ouyang, Linlin Wang, Gerard de Melo +1
Despite the pivotal role of numerical reasoning as the cornerstone of mathematical capabilities in large language models (LLMs) across applications, few benchmarks evaluate LLMs by…
cs.AI2024
How Do Humans Write Code? Large Models Do It the Same Way Too
Long Li, Xuzheng He, Haozhe Wang +2
Program-of-Thought (PoT) replaces natural language-based Chain-of-Thought (CoT) as the most popular method in Large Language Models (LLMs) mathematical reasoning tasks by utilizing…