2 papers
cs.AI2026
DREAM: Deep Research Evaluation with Agentic Metrics
Elad Ben Avraham, Changhao Li, Ron Dorfman +8
Deep Research Agents generate analyst-grade reports, yet evaluating them remains challenging due to the absence of a single ground truth and the multidimensional nature of research…
cs.AI2025
A Study on Leveraging Search and Self-Feedback for Agent Reasoning
Karthikeyan K, Michelle Yuan, Elman Mansimov +6
Recent works have demonstrated that incorporating search during inference can significantly improve reasoning capabilities of language agents. Some approaches may make use of the g…