From the 1 of 6 linked papers with an AI index.
6 papers
DeepStress: Stress-Testing Deep Search Agents
Ismael Rousseau, Geraldine Damnati, Frederic Bechet
The paper introduces DeepStress, a framework that injects controlled low‑quality evidence into the retrieval component of deep search agents to evaluate how well they handle unreli…
Trade-offs in Medical LLM Adaptation: An Empirical Study in French QA
Ikram Belmadani, Oumaima El Khettari, Carlos Ramisch +3
The development of large language models (LLMs) has led to an increased focus on their adaptation to specialized domains and languages, yet the effectiveness of domain adaptation s…
CareMedEval dataset: Evaluating Critical Appraisal and Reasoning in the Biomedical Field
Doria Bonzi, Alexandre Guiggi, Frédéric Béchet +2
Critical appraisal of scientific literature is an essential skill in the biomedical field. While large language models (LLMs) can offer promising support in this task, their reliab…
Statistical Deficiency for Task Inclusion Estimation
Loïc Fosse, Frédéric Béchet, Benoît Favre +5
Tasks are central in machine learning, as they are the most natural objects to assess the capabilities of current models. The trend is to build general models able to address any t…
Factual Knowledge in Language Models: Robustness and Anomalies under Simple Temporal Context Variations
Hichem Ammar Khodja, Frédéric Béchet, Quentin Brabant +2
This paper explores the robustness of language models (LMs) to variations in the temporal context within factual knowledge. It examines whether LMs can correctly associate a tempor…
Part-Of-Speech Sensitivity of Routers in Mixture of Experts Models
Elie Antoine, Frédéric Béchet, Philippe Langlais
This study investigates the behavior of model-integrated routers in Mixture of Experts (MoE) models, focusing on how tokens are routed based on their linguistic features, specifica…