1 paper · 1 filter
Xiaojie Gu, Sherry T. Tong, Aosong Feng +8
Evaluating the reasoning abilities of large language models (LLMs) solely from final answers can obscure failures in intermediate steps, especially in multi-hop QA benchmarks witho…