1 paper
Angelie Kraft, Judith Simon, Sonja Schimmler
Question-answering (QA) and reading comprehension (RC) benchmarks are commonly used for assessing the capabilities of large language models (LLMs) to retrieve and reproduce knowled…