activity
20212026
most citedA Systematic Review of Reproducibility Research in Natural Language Processing

82 citations · 110 across the 25 of their papers we have counts for

collaborators
Showing cs.CLShow all

21 papers · 1 filter

cs.CL2026

Validating DBpedia Triple Sets for Natural Language Generation

Mark Andrade, Simon Mille, Anya Belz +1

We present a study of the quality of individual DBpedia triples from the perspective of Natural Language Generation, and propose and evaluate an approach for collecting entity-spec…

cs.CL2026

AutoForest: Automatically Generating Forest Plots from Biomedical Studies with End-to-End Evidence Extraction and Synthesis

Massimiliano Pronesti, Angelo Miculescu, Mohsin Kapdi +8

Systematic reviews rely on forest plots to synthesise quantitative evidence across biomedical studies, but generating them remains a fragmented and labour-intensive process. Resear…

cs.CL2026

A Comparative Study of Controlled Text Generation Systems Using Level-Playing-Field Evaluation Principles

Michela Lorandi, Anya Belz

Background: Many different approaches to controlled text generation (CTG) have been proposed over recent years, but it is difficult to get a clear picture of which approach perform…

cs.CL2026

Output Composability of QLoRA PEFT Modules for Plug-and-Play Attribute-Controlled Text Generation

Michela Lorandi, Anya Belz

Parameter-efficient fine-tuning (PEFT) techniques offer task-specific fine-tuning at a fraction of the cost of full fine-tuning, but require separate fine-tuning for every new task…

cs.CL2026

Beyond Outcome Verification: Verifiable Process Reward Models for Structured Reasoning

Massimiliano Pronesti, Anya Belz, Yufang Hou

Recent work on reinforcement learning with verifiable rewards (RLVR) has shown that large language models (LLMs) can be substantially improved using outcome-level verification sign…

cs.CL2025

The QCET Taxonomy of Standard Quality Criterion Names and Definitions for the Evaluation of NLP Systems

Anya Belz, Simon Mille, Craig Thomson

Prior work has shown that two NLP evaluation experiments that report results for the same quality criterion name (e.g. Fluency) do not necessarily evaluate the same aspect of quali…