1 paper
Yulong Zhang, Li Wang, Wei Du +7
Verifying multi-step reasoning in large language models is difficult due to imprecise error localization and high token costs. Existing methods either assess entire reasoning chain…