2 citations · 6 across the 4 of their papers we have counts for
6 papers · 1 filter
On the Limitations of Reference-Free Evaluations of Generated Text
Daniel Deutsch, Rotem Dror, Dan Roth
There is significant interest in developing evaluation metrics which accurately estimate the quality of generated text without the aid of a human-written reference text, which can…
Re-Examining System-Level Correlations of Automatic Summarization Evaluation Metrics
Daniel Deutsch, Rotem Dror, Dan Roth
How reliably an automatic summarization evaluation metric replicates human judgments of summary quality is quantified by system-level correlations. We identify two ways in which th…
A Statistical Analysis of Summarization Evaluation Metrics using Resampling Methods
Daniel Deutsch, Rotem Dror, Dan Roth
The quality of a summarization evaluation metric is quantified by calculating the correlation between its scores and human annotations across a large number of summaries. Currently…
The Structured Weighted Violations MIRA
Dor Ringel, Rotem Dror, Roi Reichart
We present the Structured Weighted Violation MIRA (SWVM), a new structured prediction algorithm that is based on an hybridization between MIRA (Crammer and Singer, 2003) and the st…
Appendix - Recommended Statistical Significance Tests for NLP Tasks
Rotem Dror, Roi Reichart
Statistical significance testing plays an important role when drawing conclusions from experimental results in NLP papers. Particularly, it is a valuable tool when one would like t…
Replicability Analysis for Natural Language Processing: Testing Significance with Multiple Datasets
Rotem Dror, Gili Baumer, Marina Bogomolov +1
With the ever-growing amounts of textual data from a large variety of languages, domains, and genres, it has become standard to evaluate NLP algorithms on multiple datasets in orde…