collaborators

5 papers

cs.CL2026

Finetuning with Scientific Data Increases Hallucinations: A Multi-domain Factuality Evaluation of LLMs

Raia Abu Ahmad, Nikolas Rauscher, Ekaterina Borisova +3

Large language models (LLMs) are increasingly used to communicate and explain scientific concepts, yet their tendency to hallucinate poses significant risks in this high stakes use…

cs.CL2026

SciLaD: A Large-Scale, Transparent, Reproducible Dataset for Natural Scientific Language Processing

Luca Foppiano, Sotaro Takeshita, Pedro Ortiz Suarez +6

SciLaD is a novel, large-scale dataset of scientific language constructed entirely using open-source frameworks and publicly available data sources. It comprises a curated English…

cs.CL2025

NFDI4DS Shared Tasks for Scholarly Document Processing

Raia Abu Ahmad, Rana Abdulla, Tilahun Abedissa Taffa +18

Shared tasks are powerful tools for advancing research through community-based standardised evaluation. As such, they play a key role in promoting findable, accessible, interoperab…

cs.CL2025

Table Understanding and (Multimodal) LLMs: A Cross-Domain Case Study on Scientific vs. Non-Scientific Data

Ekaterina Borisova, Fabio Barth, Nils Feldhus +5

Tables are among the most widely used tools for representing structured data in research, business, medicine, and education. Although LLMs demonstrate strong performance in downstr…

cs.CL2025

How desirable is alignment between LLMs and linguistically diverse human users?

Pia Knoeferle, Sebastian Möller, Dorothea Kolossa +2

We discuss how desirable it is that Large Language Models (LLMs) be able to adapt or align their language behavior with users who may be diverse in their language use. User diversi…