activity
20162022
most citedEvaluating the State-of-the-Art of End-to-End Natural Language Generation: The E2E NLG Challenge

185 citations · 519 across the 16 of their papers we have counts for

collaborators
Showing cs.CLShow all

33 papers · 1 filter

cs.CL2024

NLP Verification: Towards a General Methodology for Certifying Robustness

Marco Casadio, Tanvi Dinkar, Ekaterina Komendantskaya +6

Machine Learning (ML) has exhibited substantial success in the field of Natural Language Processing (NLP). For example large language models have empirically proven to be capable o…

cs.CL2023

Understanding Counterspeech for Online Harm Mitigation

Yi-Ling Chung, Gavin Abercrombie, Florence Enock +2

Counterspeech offers direct rebuttals to hateful speech by challenging perpetrators of hate and showing support to targets of abuse. It provides a promising alternative to more con…

cs.CL20231 cited

The Dangers of trusting Stochastic Parrots: Faithfulness and Trust in Open-domain Conversational Question Answering

Sabrina Chiesurin, Dimitris Dimakopoulos, Marco Antonio Sobrevilla Cabezudo +4

Large language models are known to produce output which sounds fluent and convincing, but is also often wrong, e.g. "unfaithful" with respect to a rationale as retrieved from a kno…

cs.CL20232 cited

iLab at SemEval-2023 Task 11 Le-Wi-Di: Modelling Disagreement or Modelling Perspectives?

Nikolas Vitsakis, Amit Parekh, Tanvi Dinkar +3

There are two competing approaches for modelling annotator disagreement: distributional soft-labelling approaches (which aim to capture the level of disagreement) or modelling pers…

cs.CL2023

ANTONIO: Towards a Systematic Method of Generating NLP Benchmarks for Verification

Marco Casadio, Luca Arnaboldi, Matthew L. Daggitt +5

Verification of machine learning models used in Natural Language Processing (NLP) is known to be a hard problem. In particular, many known neural network verification methods that…

cs.CL2023

Missing Information, Unresponsive Authors, Experimental Flaws: The Impossibility of Assessing the Reproducibility of Previous Human Evaluations in NLP

Anya Belz, Craig Thomson, Ehud Reiter +39

We report our efforts in identifying a set of previous human evaluations in NLP that would be suitable for a coordinated study examining what makes human evaluations in NLP more/le…