natural language processing

DeepStress: Stress-Testing Deep Search Agents

arXiv:2607.13920

summary

The paper introduces DeepStress, a framework that injects controlled low‑quality evidence into the retrieval component of deep search agents to evaluate how well they handle unreliable, irrelevant, or untrustworthy documents, and proposes new metrics for assessing their performance on benchmarks like HotpotQA and BrowseCompPlus.

Abstract

While search agents demonstrate impressive capabilities in multi-step question answering, their robustness to poor-quality evidence remains under-explored. This phenomenon occurs rarely in realistic benchmarks but can lead to dramatic failure in real life applications. Therefore in this study we propose DeepStress, a stress testing framework that controls the frequency of challenging evidence by replacing the retrieval module of search agents with a controlled synthetic environment. We use this framework to control three dimensions that can affect document reliability: trustworthiness, relevance, and factuality. Testing several search agents on HotpotQA and BrowseCompPlus, we demonstrate that agents exhibit substantial differences in their ability to handle unreliable information and propose new metrics that better document systems outcomes as well as the interactions between conflicting parametric and retrieved knowledge.

9 pages preprint

Topics & keywords

#stress testing#search agents#multi-step question answering#evidence reliability#evaluation metrics#synthetic retrievaldeep search agentssynthetic retrieval environmenttrustworthinessrelevancefactualityHotpotQABrowseCompPlusstress testing framework
DeepStress: Stress-Testing Deep Search Agents · wovepaper