2 citations · 2 across the 16 of their papers we have counts for
Showing 2026 · cs.AIShow all
2 papers · 2 filters
cs.AI2026
FrontierChallenge: Evaluating Scientific Workflow Completion
Liangcai Su, Zhaopeng Feng, Zhuo Chen +14
Scientific agents increasingly analyze data, execute code, and produce research artifacts, yet most benchmarks emphasize final answers, isolated programs, or a single domain. We in…
cs.AI2026
MiroEval: Benchmarking Multimodal Deep Research Agents in Process and Outcome
Fangda Ye, Yuxin Hu, Pengxiang Zhu +19
Recent progress in deep research systems has been impressive, but evaluation still lags behind real user needs. Existing benchmarks predominantly assess final reports using fixed r…