5 papers
The Provenance Gap in Clinical AI: Evidence-Traceable Temporal Knowledge Graphs for Rare Disease Reasoning
Md Shamim Ahmed, Maja Dusanic, Moritz Nikolai Kirschner +4
Frontier large language models generate clinically accurate outputs, but their citations are often fabricated. We term this the Provenance Gap. We tested five frontier LLMs across…
SommBench: Assessing Sommelier Expertise of Language Models
William Brach, Tomas Bedej, Jacob Nielsen +10
With the rapid advances of large language models, it becomes increasingly important to systematically evaluate their multilingual and multicultural capabilities. Previous cultural…
FlexMoRE: A Flexible Mixture of Rank-heterogeneous Experts for Efficient Federatedly-trained Large Language Models
Annemette Brok Pirchert, Jacob Nielsen, Mogens Henrik From +2
Recent advances in mixture-of-experts architectures have shown that individual experts models can be trained federatedly, i.e., in isolation from other experts by using a common ba…
SDUs DAISY: A Benchmark for Danish Culture
Jacob Nielsen, Stine L. Beltoft, Peter Schneider-Kamp +1
We introduce a new benchmark for Danish culture via cultural heritage, Daisy, based on the curated topics from the Danish Culture Canon 2006. For each artifact in the culture canon…
Chain of Summaries: Summarization Through Iterative Questioning
William Brach, Kristián Košťál, Lukas Galke Poech
Large Language Models (LLMs) are increasingly using external web content. However, much of this content is not easily digestible by LLMs due to LLM-unfriendly formats and limitatio…