3 papers
cs.CL2025
Confabulations from ACL Publications (CAP): A Dataset for Scientific Hallucination Detection
Federica Gamba, Aman Sinha, Timothee Mickus +12
We introduce the CAP (Confabulations from ACL Publications) dataset, a multilingual resource for studying hallucinations in large language models (LLMs) within scientific text gene…
cs.CL2025
Can Out-of-Distribution Evaluations Uncover Reliance on Shortcuts? A Case Study in Question Answering
Michal Štefánik, Timothee Mickus, Marek Kadlčík +2
A majority of recent work in AI assesses models' generalization capabilities through the lens of performance on out-of-distribution (OOD) datasets. Despite their practicality, such…
cs.CL2025
Pre-trained Language Models Learn Remarkably Accurate Representations of Numbers
Marek Kadlčík, Michal Štefánik, Timothee Mickus +2
Pretrained language models (LMs) are prone to arithmetic errors. Existing work showed limited success in probing numeric values from models' representations, indicating that these…