activity
20212026
most citedA Systematic Review of Reproducibility Research in Natural Language Processing

82 citations · 99 across the 12 of their papers we have counts for

collaborators

14 papers

cs.CL2026

Beyond Outcome Verification: Verifiable Process Reward Models for Structured Reasoning

Massimiliano Pronesti, Anya Belz, Yufang Hou

Recent work on reinforcement learning with verifiable rewards (RLVR) has shown that large language models (LLMs) can be substantially improved using outcome-level verification sign…

cs.CL2025

The QCET Taxonomy of Standard Quality Criterion Names and Definitions for the Evaluation of NLP Systems

Anya Belz, Simon Mille, Craig Thomson

Prior work has shown that two NLP evaluation experiments that report results for the same quality criterion name (e.g. Fluency) do not necessarily evaluate the same aspect of quali…

cs.CL2025

Induction Signatures Are Not Enough: A Matched-Compute Study of Load-Bearing Structure in In-Context Learning

Mohammed Sabry, Anya Belz

Mechanism-targeted synthetic data is increasingly proposed as a way to steer pretraining toward desirable capabilities, but it remains unclear how such interventions should be eval…

cs.CL2025

Query-driven Document-level Scientific Evidence Extraction from Biomedical Studies

Massimiliano Pronesti, Joao Bettencourt-Silva, Paul Flanagan +4

Extracting scientific evidence from biomedical studies for clinical research questions (e.g., Does stem cell transplantation improve quality of life in patients with medically refr…

cs.CL2025

QRA++: Quantified Reproducibility Assessment for Common Types of Results in Natural Language Processing

Anya Belz

Reproduction studies reported in NLP provide individual data points which in combination indicate worryingly low levels of reproducibility in the field. Because each reproduction s…

cs.AI2025

Enhancing Study-Level Inference from Clinical Trial Papers via Reinforcement Learning-Based Numeric Reasoning

Massimiliano Pronesti, Michela Lorandi, Paul Flanagan +3

Systematic reviews in medicine play a critical role in evidence-based decision-making by aggregating findings from multiple studies. A central bottleneck in automating this process…