4 papers · 1 filter
Evaluating RAG Metrics in Applied Contexts: An Experiment, Its Findings and Its Limitations
Quentin Brabant
This paper reports an empirical study evaluating the relevance of several RAG metrics. The experiment is based on a question-answering dataset created by human annotators from busi…
Factual Knowledge in Language Models: Robustness and Anomalies under Simple Temporal Context Variations
Hichem Ammar Khodja, Frédéric Béchet, Quentin Brabant +2
This paper explores the robustness of language models (LMs) to variations in the temporal context within factual knowledge. It examines whether LMs can correctly associate a tempor…
Question Generation in Knowledge-Driven Dialog: Explainability and Evaluation
Juliette Faille, Quentin Brabant, Gwenole Lecorve +2
We explore question generation in the context of knowledge-grounded dialogs focusing on explainability and evaluation. Inspired by previous work on planning-based summarisation, we…
WikiFactDiff: A Large, Realistic, and Temporally Adaptable Dataset for Atomic Factual Knowledge Update in Causal Language Models
Hichem Ammar Khodja, Frédéric Béchet, Quentin Brabant +2
The factuality of large language model (LLMs) tends to decay over time since events posterior to their training are "unknown" to them. One way to keep models up-to-date could be fa…