6 papers
TelcoLM: collecting data, adapting, and benchmarking language models for the telecommunication domain
Camille Barboule, Viet-Phi Huynh, Adrien Bufort +3
Despite outstanding processes in many tasks, Large Language Models (LLMs) still lack accuracy when dealing with highly technical domains. Especially, telecommunications (telco) is…
Question Generation in Knowledge-Driven Dialog: Explainability and Evaluation
Juliette Faille, Quentin Brabant, Gwenole Lecorve +2
We explore question generation in the context of knowledge-grounded dialogs focusing on explainability and evaluation. Inspired by previous work on planning-based summarisation, we…
WikiFactDiff: A Large, Realistic, and Temporally Adaptable Dataset for Atomic Factual Knowledge Update in Causal Language Models
Hichem Ammar Khodja, Frédéric Béchet, Quentin Brabant +2
The factuality of large language model (LLMs) tends to decay over time since events posterior to their training are "unknown" to them. One way to keep models up-to-date could be fa…
WEBDial, a Multi-domain, Multitask Statistical Dialogue Framework with RDF
Morgan Veyret, Jean-Baptiste Duchene, Kekeli Afonouvi +3
Typically available dialogue frameworks have adopted a semantic representation based on dialogue-acts and slot-value pairs. Despite its simplicity, this representation has disadvan…
KGConv, a Conversational Corpus grounded in Wikidata
Quentin Brabant, Gwenole Lecorve, Lina M. Rojas-Barahona +1
We present KGConv, a large, conversational corpus of 71k conversations where each question-answer pair is grounded in a Wikidata fact. Conversations contain on average 8.6 question…
Age Recommendation from Texts and Sentences for Children
Rashedur Rahman, Gwénolé Lecorvé, Nicolas Béchet
Children have less text understanding capability than adults. Moreover, this capability differs among the children of different ages. Hence, automatically predicting a recommended…