natural language processing

ResearchQA: Benchmarking Citation-Grounded Question-Answering on Scientific Papers

arXiv:2607.11074

summary

The paper presents ResearchQA, a benchmark of over 6,000 question‑answer pairs from scientific papers designed to evaluate citation‑grounded answering, and uses it to compare several large language models on how well they cite supporting passages.

Abstract

Large language models are increasingly used to assist scientific reading, but existing evaluation methods often fail to detect whether answers are supported by verifiable citations. We introduce ResearchQA, a benchmark of 6,211 single-paper question-answer pairs from 494 open-access papers spanning eight domains and four question types: lookup, comprehension, multi-hop, and adversarial. ResearchQA is designed for citation-grounded evaluation: it permits multiple valid supporting passages for a claim and rewards grounded refusal when the source paper does not support an answer. We evaluate eight leading closed- and open-weight models in a citation-grounded chat-with-paper setting using a deterministic citation matcher and an LLM-based rubric evaluator. Citation-based metrics separate systems more clearly than LLM-evaluator scores: section coverage and citation accuracy vary substantially across models, while evaluator scores remain tightly compressed. We further find that open-weight models approach the best closed-model citation accuracy while achieving 3 to 6 times lower per-example latency. We release the benchmark, evaluation harness, and evaluator prompt.

19 pages, 9 figures

Topics & keywords

#question answering#citation grounding#benchmark#scientific papers#large language modelsResearchQAcitation accuracymulti-hop questionsdeterministic citation matcherLLM-based rubric evaluator
ResearchQA: Benchmarking Citation-Grounded Question-Answering on Scientific Papers · wovepaper