activity
20172025
most citedThe GEM Benchmark: Natural Language Generation, its Evaluation and Metrics

52 citations · 71 across the 10 of their papers we have counts for

collaborators

14 papers

cs.CL2025

Enhancing Long Document Long Form Summarisation with Self-Planning

Xiaotang Du, Rohit Saxena, Laura Perez-Beltrachini +2

We introduce a novel approach for long context summarisation, highlight-guided generation, that leverages sentence-level information as a content plan to improve the traceability a…

cs.CL2025

Uncertainty Quantification in Retrieval Augmented Question Answering

Laura Perez-Beltrachini, Mirella Lapata

Retrieval augmented Question Answering (QA) helps QA models overcome knowledge gaps by incorporating retrieved evidence, typically a set of passages, alongside the question at test…

cs.CL2024

Leveraging Entailment Judgements in Cross-Lingual Summarisation

Huajian Zhang, Laura Perez-Beltrachini

Synthetically created Cross-Lingual Summarisation (CLS) datasets are prone to include document-summary pairs where the reference summary is unfaithful to the corresponding document…

cs.CL2024★ 6 cited

The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models

Giwon Hong, Aryo Pradipta Gema, Rohit Saxena +8

Large Language Models (LLMs) have transformed the Natural Language Processing (NLP) landscape with their remarkable ability to understand and generate human-like text. However, the…

cs.CL2024

Fine-Grained Natural Language Inference Based Faithfulness Evaluation for Diverse Summarisation Tasks

Huajian Zhang, Yumo Xu, Laura Perez-Beltrachini

We study existing approaches to leverage off-the-shelf Natural Language Inference (NLI) models for the evaluation of summary faithfulness and argue that these are sub-optimal due t…

cs.CL2023

Improving User Controlled Table-To-Text Generation Robustness

Hanxu Hu, Yunqing Liu, Zhongyi Yu +1

In this work we study user controlled table-to-text generation where users explore the content in a table by selecting cells and reading a natural language description thereof auto…