185 citations · 519 across the 16 of their papers we have counts for
33 papers · 1 filter
NLP Verification: Towards a General Methodology for Certifying Robustness
Marco Casadio, Tanvi Dinkar, Ekaterina Komendantskaya +6
Machine Learning (ML) has exhibited substantial success in the field of Natural Language Processing (NLP). For example large language models have empirically proven to be capable o…
Understanding Counterspeech for Online Harm Mitigation
Yi-Ling Chung, Gavin Abercrombie, Florence Enock +2
Counterspeech offers direct rebuttals to hateful speech by challenging perpetrators of hate and showing support to targets of abuse. It provides a promising alternative to more con…
The Dangers of trusting Stochastic Parrots: Faithfulness and Trust in Open-domain Conversational Question Answering
Sabrina Chiesurin, Dimitris Dimakopoulos, Marco Antonio Sobrevilla Cabezudo +4
Large language models are known to produce output which sounds fluent and convincing, but is also often wrong, e.g. "unfaithful" with respect to a rationale as retrieved from a kno…
iLab at SemEval-2023 Task 11 Le-Wi-Di: Modelling Disagreement or Modelling Perspectives?
Nikolas Vitsakis, Amit Parekh, Tanvi Dinkar +3
There are two competing approaches for modelling annotator disagreement: distributional soft-labelling approaches (which aim to capture the level of disagreement) or modelling pers…
ANTONIO: Towards a Systematic Method of Generating NLP Benchmarks for Verification
Marco Casadio, Luca Arnaboldi, Matthew L. Daggitt +5
Verification of machine learning models used in Natural Language Processing (NLP) is known to be a hard problem. In particular, many known neural network verification methods that…
Missing Information, Unresponsive Authors, Experimental Flaws: The Impossibility of Assessing the Reproducibility of Previous Human Evaluations in NLP
Anya Belz, Craig Thomson, Ehud Reiter +39
We report our efforts in identifying a set of previous human evaluations in NLP that would be suitable for a coordinated study examining what makes human evaluations in NLP more/le…