5 papers · 1 filter
KD4MT: A Survey of Knowledge Distillation for Machine Translation
Ona de Gibert, Joseph Attieh, Timothee Mickus +2
Knowledge Distillation (KD) as a research area has gained a lot of traction in recent years as a compression tool to address challenges related to ever-larger models in NLP. Remark…
Confabulations from ACL Publications (CAP): A Dataset for Scientific Hallucination Detection
Federica Gamba, Aman Sinha, Timothee Mickus +12
We introduce the CAP (Confabulations from ACL Publications) dataset, a multilingual resource for studying hallucinations in large language models (LLMs) within scientific text gene…
Language Models Learn Universal Representations of Numbers and Here's Why You Should Care
Michal Štefánik, Timothee Mickus, Marek Kadlčík +7
Prior work has shown that large language models (LLMs) often converge to accurate input embedding for numbers, based on sinusoidal representations. In this work, we quantify that t…
Can Out-of-Distribution Evaluations Uncover Reliance on Shortcuts? A Case Study in Question Answering
Michal Štefánik, Timothee Mickus, Marek Kadlčík +2
A majority of recent work in AI assesses models' generalization capabilities through the lens of performance on out-of-distribution (OOD) datasets. Despite their practicality, such…
Pre-trained Language Models Learn Remarkably Accurate Representations of Numbers
Marek Kadlčík, Michal Štefánik, Timothee Mickus +2
Pretrained language models (LMs) are prone to arithmetic errors. Existing work showed limited success in probing numeric values from models' representations, indicating that these…