activity
20202022
most citedRe-Examining System-Level Correlations of Automatic Summarization Evaluation Metrics

2 citations · 5 across the 5 of their papers we have counts for

collaborators
Showing cs.CLShow all

8 papers · 1 filter

cs.CL20222 cited

On the Limitations of Reference-Free Evaluations of Generated Text

Daniel Deutsch, Rotem Dror, Dan Roth

There is significant interest in developing evaluation metrics which accurately estimate the quality of generated text without the aid of a human-written reference text, which can…

cs.CL20221 cited

Repro: An Open-Source Library for Improving the Reproducibility and Usability of Publicly Available Research Code

Daniel Deutsch, Dan Roth

We introduce Repro, an open-source library which aims at improving the reproducibility and usability of research code. The library provides a lightweight Python API for running sof…

cs.CL20222 cited

Re-Examining System-Level Correlations of Automatic Summarization Evaluation Metrics

Daniel Deutsch, Rotem Dror, Dan Roth

How reliably an automatic summarization evaluation metric replicates human judgments of summary quality is quantified by system-level correlations. We identify two ways in which th…

cs.CL2022

Benchmarking Answer Verification Methods for Question Answering-Based Summarization Evaluation Metrics

Daniel Deutsch, Dan Roth

Question answering-based summarization evaluation metrics must automatically determine whether the QA model's prediction is correct or not, a task known as answer verification. In…

cs.CL2021

A Statistical Analysis of Summarization Evaluation Metrics using Resampling Methods

Daniel Deutsch, Rotem Dror, Dan Roth

The quality of a summarization evaluation metric is quantified by calculating the correlation between its scores and human annotations across a large number of summaries. Currently…

cs.CL2020

Understanding the Extent to which Summarization Evaluation Metrics Measure the Information Quality of Summaries

Daniel Deutsch, Dan Roth

Reference-based metrics such as ROUGE or BERTScore evaluate the content quality of a summary by comparing the summary to a reference. Ideally, this comparison should measure the su…