activity
20212026
most citedA Systematic Review of Reproducibility Research in Natural Language Processing

82 citations · 110 across the 25 of their papers we have counts for

collaborators
Showing 2025Show all

5 papers · 1 filter

cs.CL2025

The QCET Taxonomy of Standard Quality Criterion Names and Definitions for the Evaluation of NLP Systems

Anya Belz, Simon Mille, Craig Thomson

Prior work has shown that two NLP evaluation experiments that report results for the same quality criterion name (e.g. Fluency) do not necessarily evaluate the same aspect of quali…

cs.CL2025

Induction Signatures Are Not Enough: A Matched-Compute Study of Load-Bearing Structure in In-Context Learning

Mohammed Sabry, Anya Belz

Mechanism-targeted synthetic data is increasingly proposed as a way to steer pretraining toward desirable capabilities, but it remains unclear how such interventions should be eval…

cs.CL2025

Query-driven Document-level Scientific Evidence Extraction from Biomedical Studies

Massimiliano Pronesti, Joao Bettencourt-Silva, Paul Flanagan +4

Extracting scientific evidence from biomedical studies for clinical research questions (e.g., Does stem cell transplantation improve quality of life in patients with medically refr…

cs.CL2025

QRA++: Quantified Reproducibility Assessment for Common Types of Results in Natural Language Processing

Anya Belz

Reproduction studies reported in NLP provide individual data points which in combination indicate worryingly low levels of reproducibility in the field. Because each reproduction s…

cs.AI2025

Enhancing Study-Level Inference from Clinical Trial Papers via Reinforcement Learning-Based Numeric Reasoning

Massimiliano Pronesti, Michela Lorandi, Paul Flanagan +3

Systematic reviews in medicine play a critical role in evidence-based decision-making by aggregating findings from multiple studies. A central bottleneck in automating this process…