4 papers
Averaging Bias: Human Faithfulness Annotations are not Locally Faithful
Huajian Zhang, Yiyang Feng, Jiawei Zhou
Evaluation of faithfulness of text summarization treats a model generated summary as faithful only if every of its sentences is supported by the source document: a strict conjuncti…
Self-Improvement of Large Language Models: A Technical Overview and Future Outlook
Haoyan Yang, Mario Xerri, Solha Park +4
As large language models (LLMs) continue to advance, improving them solely through human supervision is becoming increasingly costly and limited in scalability. As models approach…
Leveraging Entailment Judgements in Cross-Lingual Summarisation
Huajian Zhang, Laura Perez-Beltrachini
Synthetically created Cross-Lingual Summarisation (CLS) datasets are prone to include document-summary pairs where the reference summary is unfaithful to the corresponding document…
Fine-Grained Natural Language Inference Based Faithfulness Evaluation for Diverse Summarisation Tasks
Huajian Zhang, Yumo Xu, Laura Perez-Beltrachini
We study existing approaches to leverage off-the-shelf Natural Language Inference (NLI) models for the evaluation of summary faithfulness and argue that these are sub-optimal due t…