30 citations · 131 across the 21 of their papers we have counts for
24 papers · 1 filter
Layer or Representation Space: What makes BERT-based Evaluation Metrics Robust?
Doan Nam Long Vu, Nafise Sadat Moosavi, Steffen Eger
The evaluation of recent embedding-based evaluation metrics for text generation is primarily based on measuring their correlation with human evaluations on standard benchmarks. How…
Towards Explainable Evaluation Metrics for Natural Language Generation
Christoph Leiter, Piyawat Lertvittayakumjorn, Marina Fomicheva +3
Unlike classical lexical overlap metrics such as BLEU, most current evaluation metrics (such as BERTScore or MoverScore) are based on black-box language models such as BERT or XLM-…
Global Explainability of BERT-Based Evaluation Metrics by Disentangling along Linguistic Factors
Marvin Kaster, Wei Zhao, Steffen Eger
Evaluation metrics are a key ingredient for progress of text generation systems. In recent years, several BERT-based evaluation metrics have been proposed (including BERTScore, Mov…
Better than Average: Paired Evaluation of NLP Systems
Maxime Peyrard, Wei Zhao, Steffen Eger +1
Evaluation in NLP is usually done by comparing the scores of competing systems independently averaged over a common set of test instances. In this work, we question the use of aver…
The Eval4NLP Shared Task on Explainable Quality Estimation: Overview and Results
Marina Fomicheva, Piyawat Lertvittayakumjorn, Wei Zhao +2
In this paper, we introduce the Eval4NLP-2021shared task on explainable quality estimation. Given a source-translation pair, this shared task requires not only to provide a sentenc…
Diachronic Analysis of German Parliamentary Proceedings: Ideological Shifts through the Lens of Political Biases
Tobias Walter, Celina Kirschner, Steffen Eger +3
We analyze bias in historical corpora as encoded in diachronic distributional semantic models by focusing on two specific forms of bias, namely a political (i.e., anti-communism) a…