2 papers
cs.CL2024
Knowledge Distillation vs. Pretraining from Scratch under a Fixed (Computation) Budget
Minh Duc Bui, Fabian David Schmidt, Goran Glavaš +1
Compared to standard language model (LM) pretraining (i.e., from scratch), Knowledge Distillation (KD) entails an additional forward pass through a teacher model that is typically…
cs.CL2022
Geographic Adaptation of Pretrained Language Models
Valentin Hofmann, Goran Glavaš, Nikola Ljubešić +2
While pretrained language models (PLMs) have been shown to possess a plethora of linguistic knowledge, the existing body of research has largely neglected extralinguistic knowledge…