6 papers · 1 filter
Language Models Learn Universal Representations of Numbers and Here's Why You Should Care
Michal Å tefánik, Timothee Mickus, Marek KadlÄÃk +7
Prior work has shown that large language models (LLMs) often converge to accurate input embedding for numbers, based on sinusoidal representations. In this work, we quantify that t…
Life Cycle-Aware Evaluation of Knowledge Distillation for Machine Translation: Environmental Impact and Translation Quality Trade-offs
Joseph Attieh, Timothee Mickus, Anne-Laure Ligozat +2
Knowledge distillation (KD) is a tool to compress a larger system (teacher) into a smaller one (student). In machine translation, studies typically report only the translation qual…
KD4MT: A Survey of Knowledge Distillation for Machine Translation
Ona de Gibert, Joseph Attieh, Timothee Mickus +2
Knowledge Distillation (KD) as a research area has gained a lot of traction in recent years as a compression tool to address challenges related to ever-larger models in NLP. Remark…
Confabulations from ACL Publications (CAP): A Dataset for Scientific Hallucination Detection
Federica Gamba, Aman Sinha, Timothee Mickus +12
We introduce the CAP (Confabulations from ACL Publications) dataset, a multilingual resource for studying hallucinations in large language models (LLMs) within scientific text gene…
Pre-trained Language Models Learn Remarkably Accurate Representations of Numbers
Marek KadlÄÃk, Michal Å tefánik, Timothee Mickus +2
Pretrained language models (LMs) are prone to arithmetic errors. Existing work showed limited success in probing numeric values from models' representations, indicating that these…
Can Out-of-Distribution Evaluations Uncover Reliance on Shortcuts? A Case Study in Question Answering
Michal Å tefánik, Timothee Mickus, Marek KadlÄÃk +2
A majority of recent work in AI assesses models' generalization capabilities through the lens of performance on out-of-distribution (OOD) datasets. Despite their practicality, such…