activity
20202026
most citedSUPERT: Towards New Frontiers in Unsupervised Evaluation Metrics for Multi-Document Summarization

14 citations · 37 across the 9 of their papers we have counts for

collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2026

How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures

Paul Osemudiame Oamen, Owusu-Banahene Osei, Ananya Mukherjee +4

Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what they see in an image), with limited attent…

cs.CL2021

Global Explainability of BERT-Based Evaluation Metrics by Disentangling along Linguistic Factors

Marvin Kaster, Wei Zhao, Steffen Eger

Evaluation metrics are a key ingredient for progress of text generation systems. In recent years, several BERT-based evaluation metrics have been proposed (including BERTScore, Mov…

cs.CL202114 cited

Better than Average: Paired Evaluation of NLP Systems

Maxime Peyrard, Wei Zhao, Steffen Eger +1

Evaluation in NLP is usually done by comparing the scores of competing systems independently averaged over a common set of test instances. In this work, we question the use of aver…

cs.CL20211 cited

The Eval4NLP Shared Task on Explainable Quality Estimation: Overview and Results

Marina Fomicheva, Piyawat Lertvittayakumjorn, Wei Zhao +2

In this paper, we introduce the Eval4NLP-2021shared task on explainable quality estimation. Given a source-translation pair, this shared task requires not only to provide a sentenc…

cs.CL20204 cited

On the Limitations of Cross-lingual Encoders as Exposed by Reference-Free Machine Translation Evaluation

Wei Zhao, Goran Glavaš, Maxime Peyrard +3

Evaluation of cross-lingual encoders is usually performed either via zero-shot cross-lingual transfer in supervised downstream tasks or via unsupervised cross-lingual textual simil…

cs.CL202014 cited

SUPERT: Towards New Frontiers in Unsupervised Evaluation Metrics for Multi-Document Summarization

Yang Gao, Wei Zhao, Steffen Eger

We study unsupervised multi-document summarization evaluation metrics, which require neither human-written reference summaries nor human annotations (e.g. preferences, ratings, etc…