2 papers
cs.AI2026
DRBENCHER: Can Your Agent Identify the Entity, Retrieve Its Properties and Do the Math?
Young-Suk Lee, Ramon Fernandez Astudillo, Radu Florian
Deep research agents increasingly interleave web browsing with multi-step computation, yet existing benchmarks evaluate these capabilities in isolation, creating a blind spot in as…
cs.CL2025
MTRAG: A Multi-Turn Conversational Benchmark for Evaluating Retrieval-Augmented Generation Systems
Yannis Katsis, Sara Rosenthal, Kshitij Fadnis +7
Retrieval-augmented generation (RAG) has recently become a very popular task for Large Language Models (LLMs). Evaluating them on multi-turn RAG conversations, where the system is…