1 paper
Chulin Xie, Yangsibo Huang, Chiyuan Zhang +6
Large language models (LLMs) achieve good performance on challenging reasoning benchmarks, yet could also make basic reasoning mistakes. This contrasting behavior is puzzling when…