evaluation metrics 1evidence reliability 1multi-step question answering 1search agents 1stress testing 1synthetic retrieval 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.CL2026
DeepStress: Stress-Testing Deep Search Agents
Ismael Rousseau, Geraldine Damnati, Frederic Bechet
The paper introduces DeepStress, a framework that injects controlled low‑quality evidence into the retrieval component of deep search agents to evaluate how well they handle unreli…
cs.CL2025
O_FT@EvalLLM2025 : étude comparative de choix de données et de stratégies d'apprentissage pour l'adaptation de modèles de langue à un domaine
Ismaël Rousseau, Claire Perroux, Pierre Adam +5
This paper presents the work carried out by the O_FT team, joint with Orange and Ouest-France, on adapting language models to the defense domain as part of the EvalLLM2025 challeng…