4 papers
Knowledge Index of Noah's Ark
Sheng Jin, Minghao Liu, Yunze Xiao +24
Knowledge benchmarks for LLMs face three issues: scaling-driven designs that do not operationalize disciplinary representativeness; flat-payment annotation that permits lazy consen…
CriticLean: Critic-Guided Reinforcement Learning for Mathematical Formalization
Zhongyuan Peng, Yifan Yao, Kaijing Ma +16
Translating natural language mathematical statements into formal, executable code is a fundamental challenge in automated theorem proving. While prior work has focused on generatio…
FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models
Zhouliang Yu, Ruotian Peng, Keyi Ding +10
Formal mathematical reasoning remains a critical challenge for artificial intelligence, hindered by limitations of existing benchmarks in scope and scale. To address this, we prese…
SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines
P Team, Xinrun Du, Yifan Yao +94
Large language models (LLMs) have demonstrated remarkable proficiency in mainstream academic disciplines such as mathematics, physics, and computer science. However, human knowledg…