1 paper
Yunchong Huang, Gianni Barlacchi, Sandro Pezzelle
Large language models (LLMs) perform well on well-posed questions, yet standard question-answering (QA) benchmarks remain far from solved. We argue that this gap is partly due to u…