activity
20172026
most citedCombining Modular Skills in Multitask Learning

17 citations · 42 across the 38 of their papers we have counts for

collaborators
Showing 2023Show all

6 papers · 1 filter

cs.CL2023

Are Large Language Models Temporally Grounded?

Yifu Qiu, Zheng Zhao, Yftah Ziser +3

Are Large language models (LLMs) temporally grounded? Since LLMs cannot perceive and interact with the environment, it is impossible to answer this question directly. Instead, we p…

cs.LG20232 cited

Model Merging by Uncertainty-Based Gradient Matching

Nico Daheim, Thomas Möllenhoff, Edoardo Maria Ponti +2

Models trained on different datasets can be merged by a weighted-averaging of their parameters, but why does it work and when can it fail? Here, we connect the inaccuracy of weight…

cs.CL2023

Distilling Efficient Language-Specific Models for Cross-Lingual Transfer

Alan Ansell, Edoardo Maria Ponti, Anna Korhonen +1

Massively multilingual Transformers (MMTs), such as mBERT and XLM-R, are widely used for cross-lingual transfer learning. While these are pretrained to represent hundreds of langua…

cs.CL2023

Detecting and Mitigating Hallucinations in Multilingual Summarisation

Yifu Qiu, Yftah Ziser, Anna Korhonen +2

Hallucinations pose a significant challenge to the reliability of neural models for abstractive summarisation. While automatically generated summaries may be fluent, they often lac…

cs.CL2023

Elastic Weight Removal for Faithful and Abstractive Dialogue Generation

Nico Daheim, Nouha Dziri, Mrinmaya Sachan +2

Ideally, dialogue systems should generate responses that are faithful to the knowledge contained in relevant documents. However, many models generate hallucinated responses instead…

cs.LG2023

Modular Deep Learning

Jonas Pfeiffer, Sebastian Ruder, Ivan Vulić +1

Transfer learning has recently become the dominant paradigm of machine learning. Pre-trained models fine-tuned for downstream tasks achieve better performance with fewer labelled e…