From the 1 of 11 linked papers with an AI index.
1 citations · 1 across the 4 of their papers we have counts for
14 papers · 1 filter
AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification
Lingkai Kong, Zijian Wu, Yuzhe Gu +10
The paper introduces AdvancedMathBench, a benchmark suite for evaluating large language models on generating and verifying advanced mathematical proofs, and provides an automatic v…
Lean Workbook: A large-scale Lean problem set formalized from natural language math problems
Huaiyuan Ying, Zijian Wu, Yihan Geng +3
Large language models have demonstrated impressive capabilities across various natural language processing tasks, especially in solving mathematical problems. However, large langua…
Mask-DPO: Generalizable Fine-grained Factuality Alignment of LLMs
Yuzhe Gu, Wenwei Zhang, Chengqi Lyu +2
Large language models (LLMs) exhibit hallucinations (i.e., unfaithful or nonsensical information) when serving as AI assistants in various domains. Since hallucinations always come…
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
Chengqi Lyu, Songyang Gao, Yuzhe Gu +14
Reasoning abilities, especially those for solving complex math problems, are crucial components of general intelligence. Recent advances by proprietary companies, such as o-series…
ANAH-v2: Scaling Analytical Hallucination Annotation of Large Language Models
Yuzhe Gu, Ziwei Ji, Wenwei Zhang +3
Large language models (LLMs) exhibit hallucinations in long-form question-answering tasks across various domains and wide applications. Current hallucination detection and mitigati…
Training Language Models to Critique With Multi-agent Feedback
Tian Lan, Wenwei Zhang, Chengqi Lyu +6
Critique ability, a meta-cognitive capability of humans, presents significant challenges for LLMs to improve. Recent works primarily rely on supervised fine-tuning (SFT) using crit…