1 paper · 1 filter
Jiaqi Chen, Bang Zhang, Ruotian Ma +5
Evaluating the step-by-step reliability of large language model (LLM) reasoning, such as Chain-of-Thought, remains challenging due to the difficulty and cost of obtaining high-qual…