1 citations · 1 across the 10 of their papers we have counts for
7 papers · 1 filter
FaithSieve: Fine-Grained Evaluation of Math Proofs with Faithful Formal Evidence
Ziyu Wang, Qiming Dai, Yishan Wu +1
Large language models can now generate complex, multi-step mathematical proofs, but reliably determining their correctness and localizing early logical errors remains a critical ch…
ReasFlow: Assisting Reasoning-Centric Scientific Discovery in Applied Mathematics via a Knowledge-Based Multi-Agent System
Yutong He, Daibo Li, Guohong Li +15
Recent advances in Large Language Models have fueled autonomous AI agents capable of tackling complex scientific tasks, yet existing automated research systems remain predominantly…
MECA: A Mechanism-Centered Agent for Constructing Well-Specified and Valuable Mathematical Conjectures
Wentao Long, Yunfei Zhang, Chenyi Li +1
Automatically constructing well-specified and valuable mathematical conjectures remains a central challenge in AI-assisted mathematical discovery. Many existing open problems and c…
CAM-Bench: A Benchmark for Computational and Applied Mathematics in Lean
Wentao Long, Yunfei Zhang, Chenyi Li +3
Formal theorem-proving benchmarks enable mechanically verifiable evaluation of mathematical reasoning in large language models. However, existing benchmarks mainly focus on Olympia…
MM-OptBench: A Solver-Grounded Benchmark for Multimodal Optimization Modeling
Zhong Li, Qi Huang, Yuxuan Zhu +6
Optimization modeling translates real decision-making problems into mathematical optimization models and solver-executable implementations. Although language models are increasingl…
SITA: A Framework for Structure-to-Instance Theorem Autoformalization
Chenyi Li, Wanli Ma, Zichen Wang +1
While large language models (LLMs) have shown progress in mathematical reasoning, they still face challenges in formalizing theorems that arise from instantiating abstract structur…