works on

From the 1 of 11 linked papers with an AI index.

activity
20242026
most citedMinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale

1 citations · 1 across the 4 of their papers we have counts for

collaborators
Showing cs.CLShow all

14 papers · 1 filter

cs.CL2026

AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification

Lingkai Kong, Zijian Wu, Yuzhe Gu +10

The paper introduces AdvancedMathBench, a benchmark suite for evaluating large language models on generating and verifying advanced mathematical proofs, and provides an automatic v…

cs.CL2025

Lean Workbook: A large-scale Lean problem set formalized from natural language math problems

Huaiyuan Ying, Zijian Wu, Yihan Geng +3

Large language models have demonstrated impressive capabilities across various natural language processing tasks, especially in solving mathematical problems. However, large langua…

cs.CL2025

Mask-DPO: Generalizable Fine-grained Factuality Alignment of LLMs

Yuzhe Gu, Wenwei Zhang, Chengqi Lyu +2

Large language models (LLMs) exhibit hallucinations (i.e., unfaithful or nonsensical information) when serving as AI assistants in various domains. Since hallucinations always come…

cs.CL2025

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Chengqi Lyu, Songyang Gao, Yuzhe Gu +14

Reasoning abilities, especially those for solving complex math problems, are crucial components of general intelligence. Recent advances by proprietary companies, such as o-series…

cs.CL2024

ANAH-v2: Scaling Analytical Hallucination Annotation of Large Language Models

Yuzhe Gu, Ziwei Ji, Wenwei Zhang +3

Large language models (LLMs) exhibit hallucinations in long-form question-answering tasks across various domains and wide applications. Current hallucination detection and mitigati…

cs.CL2024

Training Language Models to Critique With Multi-agent Feedback

Tian Lan, Wenwei Zhang, Chengqi Lyu +6

Critique ability, a meta-cognitive capability of humans, presents significant challenges for LLMs to improve. Recent works primarily rely on supervised fine-tuning (SFT) using crit…