1 paper
Jaeyong Ko, Pilsung Kang, Yukyung Lee
Large language models (LLMs) reach high accuracy in mathematical reasoning, but individual traces on the same problem diverge; some arrive at the correct answer while others fail.…