17 citations · 42 across the 38 of their papers we have counts for
6 papers · 1 filter
Are Large Language Models Temporally Grounded?
Yifu Qiu, Zheng Zhao, Yftah Ziser +3
Are Large language models (LLMs) temporally grounded? Since LLMs cannot perceive and interact with the environment, it is impossible to answer this question directly. Instead, we p…
Model Merging by Uncertainty-Based Gradient Matching
Nico Daheim, Thomas Möllenhoff, Edoardo Maria Ponti +2
Models trained on different datasets can be merged by a weighted-averaging of their parameters, but why does it work and when can it fail? Here, we connect the inaccuracy of weight…
Distilling Efficient Language-Specific Models for Cross-Lingual Transfer
Alan Ansell, Edoardo Maria Ponti, Anna Korhonen +1
Massively multilingual Transformers (MMTs), such as mBERT and XLM-R, are widely used for cross-lingual transfer learning. While these are pretrained to represent hundreds of langua…
Detecting and Mitigating Hallucinations in Multilingual Summarisation
Yifu Qiu, Yftah Ziser, Anna Korhonen +2
Hallucinations pose a significant challenge to the reliability of neural models for abstractive summarisation. While automatically generated summaries may be fluent, they often lac…
Elastic Weight Removal for Faithful and Abstractive Dialogue Generation
Nico Daheim, Nouha Dziri, Mrinmaya Sachan +2
Ideally, dialogue systems should generate responses that are faithful to the knowledge contained in relevant documents. However, many models generate hallucinated responses instead…
Modular Deep Learning
Jonas Pfeiffer, Sebastian Ruder, Ivan Vulić +1
Transfer learning has recently become the dominant paradigm of machine learning. Pre-trained models fine-tuned for downstream tasks achieve better performance with fewer labelled e…