4 papers
WANDR: A Benchmark for Wide and Deep Research
Vitaliy Polshkov, Marcin Pitera, Jeremy Yang +7
WANDR (Wide ANd Deep Research) is a benchmark of 500 realistic, challenging data-collection tasks for research agents. Each task requires a system to discover a large set of entiti…
DRACO: a Cross-Domain Benchmark for Deep Research Accuracy, Completeness, and Objectivity
Joey Zhong, Hao Zhang, Clare Southern +7
We present DRACO (Deep Research Accuracy, Completeness, and Objectivity), a benchmark of complex deep research tasks. These tasks, which span 10 domains and draw on information sou…
The Adoption and Usage of AI Agents: Early Evidence from Perplexity
Jeremy Yang, Noah Yonack, Kate Zyskowski +3
This paper presents the first large-scale field study of the adoption, usage intensity, and use cases of general-purpose AI agents operating in open-world web environments. Our ana…
BrowseSafe: Understanding and Preventing Prompt Injection Within AI Browser Agents
Kaiyuan Zhang, Mark Tenenholtz, Kyle Polley +3
The integration of artificial intelligence (AI) agents into web browsers introduces security challenges that go beyond traditional web application threat models. Prior work has ide…