6 papers
Knowledgeless Language Models: Suppressing Parametric Recall for Evidence-Grounded Language Modeling
Roi Cohen, Yvan Carré, Nick Lechtenbörger +5
Language models encode substantial factual knowledge in their parameters, which can lead to unreliable behavior when this knowledge is outdated, incomplete, or misaligned with the…
From Global to Local: Learning Context-Aware Graph Representations for Document Classification and Summarization
Ruangrin Ldallitsakool, Margarita Bugueño, Gerard de Melo
Recent NLP systems commonly represent documents as linear token sequences. Although this captures sequential order, it can hinder modeling long-range dependencies and global docume…
Bundesrecht: An Open Library and Corpus for German Statutory Reference Processing
Harshil Darji, Martin Heckelmann, Christina Kratsch +1
Statutory references are central to legal language understanding, but are difficult to process automatically, as they appear in compact and variable surface forms, may combine mult…
ReFACT: A Benchmark for Scientific Confabulation Detection with Positional Error Annotations
Yindong Wang, Martin PreiÃ, Margarita Bugueño +4
The mechanisms underlying scientific confabulation in Large Language Models (LLMs) remain poorly understood. We introduce ReFACT (Reddit False And Correct Texts), a benchmark of 1,…
Pretrained LLMs Learn Multiple Types of Uncertainty
Roi Cohen, Omri Fahn, Gerard de Melo
Large Language Models are known to capture real-world knowledge, allowing them to excel in many downstream tasks. Despite recent advances, these models are still prone to what are…
InFact: Informativeness Alignment for Improved LLM Factuality
Roi Cohen, Russa Biswas, Gerard de Melo
Factual completeness is a general term that captures how detailed and informative a factually correct text is. For instance, the factual sentence ``Barack Obama was born in the Uni…