From the 1 of 6 linked papers with an AI index.
5 papers · 1 filter
DeepStress: Stress-Testing Deep Search Agents
Ismael Rousseau, Geraldine Damnati, Frederic Bechet
The paper introduces DeepStress, a framework that injects controlled low‑quality evidence into the retrieval component of deep search agents to evaluate how well they handle unreli…
Trade-offs in Medical LLM Adaptation: An Empirical Study in French QA
Ikram Belmadani, Oumaima El Khettari, Carlos Ramisch +3
The development of large language models (LLMs) has led to an increased focus on their adaptation to specialized domains and languages, yet the effectiveness of domain adaptation s…
CareMedEval dataset: Evaluating Critical Appraisal and Reasoning in the Biomedical Field
Doria Bonzi, Alexandre Guiggi, Frédéric Béchet +2
Critical appraisal of scientific literature is an essential skill in the biomedical field. While large language models (LLMs) can offer promising support in this task, their reliab…
Factual Knowledge in Language Models: Robustness and Anomalies under Simple Temporal Context Variations
Hichem Ammar Khodja, Frédéric Béchet, Quentin Brabant +2
This paper explores the robustness of language models (LMs) to variations in the temporal context within factual knowledge. It examines whether LMs can correctly associate a tempor…
Part-Of-Speech Sensitivity of Routers in Mixture of Experts Models
Elie Antoine, Frédéric Béchet, Philippe Langlais
This study investigates the behavior of model-integrated routers in Mixture of Experts (MoE) models, focusing on how tokens are routed based on their linguistic features, specifica…