83 citations · 105 across the 7 of their papers we have counts for
8 papers · 1 filter
Input Matters: Evaluating Input Structure's Impact on LLM Summaries of Sports Play-by-Play
Barkavi Sundararajan, Somayajulu Sripada, Ehud Reiter
A major concern when deploying LLMs in accuracy-critical domains such as sports reporting is that the generated text may not faithfully reflect the input data. We quantify how inpu…
Consultation Checklists: Standardising the Human Evaluation of Medical Note Generation
Aleksandar Savkov, Francesco Moramarco, Alex Papadopoulos Korfiatis +3
Evaluating automatically generated text is generally hard due to the inherently subjective nature of many aspects of the output quality. This difficulty is compounded in automatic…
Human Evaluation and Correlation with Automatic Metrics in Consultation Note Generation
Francesco Moramarco, Alex Papadopoulos Korfiatis, Mark Perera +5
In recent years, machine learning models have rapidly become better at generating clinical consultation notes; yet, there is little work on how to properly evaluate the generated c…
Generation Challenges: Results of the Accuracy Evaluation Shared Task
Craig Thomson, Ehud Reiter
The Shared Task on Evaluating Accuracy focused on techniques (both manual and automatic) for evaluating the factual accuracy of texts produced by neural NLG systems, in a sports-re…
Towards objectively evaluating the quality of generated medical summaries
Francesco Moramarco, Damir Juric, Aleksandar Savkov +1
We propose a method for evaluating the quality of generated text by asking evaluators to count facts, and computing precision, recall, f-score, and accuracy from the raw counts. We…
A preliminary study on evaluating Consultation Notes with Post-Editing
Francesco Moramarco, Alex Papadopoulos Korfiatis, Aleksandar Savkov +1
Automatic summarisation has the potential to aid physicians in streamlining clerical tasks such as note taking. But it is notoriously difficult to evaluate these systems and demons…