1 paper
Biao Xiang, Soyeon Caren Han, Yihao Ding
Multi-hop question answering (QA) is widely used to evaluate the reasoning capabilities of large language models, yet most benchmarks focus on final answer correctness and overlook…