Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
MiroEval: Benchmarking Multimodal Deep Research Agents in Process and Outcome
Fangda Ye, Yuxin Hu, Pengxiang Zhu +19
Recent progress in deep research systems has been impressive, but evaluation still lags behind real user needs. Existing benchmarks predominantly assess final reports using fixed r…
cs.AI2025
First Try Matters: Revisiting the Role of Reflection in Reasoning Models
Liwei Kang, Yue Deng, Yao Xiao +3
Large language models have recently demonstrated significant gains in reasoning ability, often attributed to their capacity to generate longer chains of thought and engage in refle…