From the 1 of 6 linked papers with an AI index.
3 papers · 1 filter
DeepStress: Stress-Testing Deep Search Agents
Ismael Rousseau, Geraldine Damnati, Frederic Bechet
The paper introduces DeepStress, a framework that injects controlled low‑quality evidence into the retrieval component of deep search agents to evaluate how well they handle unreli…
Trade-offs in Medical LLM Adaptation: An Empirical Study in French QA
Ikram Belmadani, Oumaima El Khettari, Carlos Ramisch +3
The development of large language models (LLMs) has led to an increased focus on their adaptation to specialized domains and languages, yet the effectiveness of domain adaptation s…
CareMedEval dataset: Evaluating Critical Appraisal and Reasoning in the Biomedical Field
Doria Bonzi, Alexandre Guiggi, Frédéric Béchet +2
Critical appraisal of scientific literature is an essential skill in the biomedical field. While large language models (LLMs) can offer promising support in this task, their reliab…