From the 1 of 13 linked papers with an AI index.
13 papers
FaithformBench: Benchmarking Faithfulness of Mathematical Chain-of-Thought Autoformalisation
Rob Cornish, Iacopo Ghinassi, Po-Hung Yeh +7
Autoformalisation (AF) systems map natural language reasoning steps into formal statements in a proof assistant such as Lean. We consider how to assess the faithfulness of these sy…
AoA: Theorem Proving Agent over Abstract Syntax Tree of Redesigned Language
Qiyuan Xu, Joshua Ong Jun Leang, Renxi Wang +4
Interactive theorem proving (ITP) underpins program verification and formalized mathematics, but its manual effort limits scalability. LLM-based proof agents promise to ease this e…
PiCSAR: Probabilistic Confidence Selection And Ranking for Reasoning Chains
Joshua Ong Jun Leang, Zheng Zhao, Aryo Pradipta Gema +7
The paper proposes PiCSAR, a training-free scoring method that uses the joint log-likelihood of reasoning steps and final answer to select the most reliable reasoning chain from mu…
Pythagoras-Prover: Advancing Efficient Formal Proving via Augmented Lean Formalisation
Joshua Ong Jun Leang, Zheng Zhao, Mihaela CÄtÄlina Stoian +5
Modern Lean theorem provers achieve strong performance only with substantial training and inference compute, driven in part by scarce verified proof data and the long reasoning tra…
Specialize Roles, Mix Deployments: Pushing the Cost-Accuracy Frontier of LLM Agent Teams
Yinsicheng Jiang, Liang Cheng, Yeqi Huang +6
LLM agents are increasingly deployed as multi-role teams, where tasks are divided across specialized roles such as planner, executor, and verifier. In these systems, cost and accur…
Can I Have Your Order? Monte-Carlo Tree Search for Slot Filling Ordering in Diffusion Language Models
Joshua Ong Jun Leang, Yu Zhao, Mihaela CÄtÄlina Stoian +3
While plan-and-infill decoding in Masked Diffusion Models (MDMs) shows promise for mathematical and code reasoning, performance remains highly sensitive to slot infilling order, of…