Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
MiroEval: Benchmarking Multimodal Deep Research Agents in Process and Outcome
Fangda Ye, Yuxin Hu, Pengxiang Zhu +19
Recent progress in deep research systems has been impressive, but evaluation still lags behind real user needs. Existing benchmarks predominantly assess final reports using fixed r…
cs.AI2025
EcomBench: Towards Holistic Evaluation of Foundation Agents in E-commerce
Rui Min, Zile Qiao, Ze Xu +18
Foundation agents have rapidly advanced in their ability to reason and interact with real environments, making the evaluation of their core capabilities increasingly important. Whi…