activity
20172022
most citedReplicability Analysis for Natural Language Processing: Testing Significance with Multiple Datasets

2 citations · 6 across the 4 of their papers we have counts for

collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL20222 cited

On the Limitations of Reference-Free Evaluations of Generated Text

Daniel Deutsch, Rotem Dror, Dan Roth

There is significant interest in developing evaluation metrics which accurately estimate the quality of generated text without the aid of a human-written reference text, which can…

cs.CL20222 cited

Re-Examining System-Level Correlations of Automatic Summarization Evaluation Metrics

Daniel Deutsch, Rotem Dror, Dan Roth

How reliably an automatic summarization evaluation metric replicates human judgments of summary quality is quantified by system-level correlations. We identify two ways in which th…

cs.CL2021

A Statistical Analysis of Summarization Evaluation Metrics using Resampling Methods

Daniel Deutsch, Rotem Dror, Dan Roth

The quality of a summarization evaluation metric is quantified by calculating the correlation between its scores and human annotations across a large number of summaries. Currently…

cs.CL2020

The Structured Weighted Violations MIRA

Dor Ringel, Rotem Dror, Roi Reichart

We present the Structured Weighted Violation MIRA (SWVM), a new structured prediction algorithm that is based on an hybridization between MIRA (Crammer and Singer, 2003) and the st…

cs.CL2018

Appendix - Recommended Statistical Significance Tests for NLP Tasks

Rotem Dror, Roi Reichart

Statistical significance testing plays an important role when drawing conclusions from experimental results in NLP papers. Particularly, it is a valuable tool when one would like t…

cs.CL20172 cited

Replicability Analysis for Natural Language Processing: Testing Significance with Multiple Datasets

Rotem Dror, Gili Baumer, Marina Bogomolov +1

With the ever-growing amounts of textual data from a large variety of languages, domains, and genres, it has become standard to evaluate NLP algorithms on multiple datasets in orde…