Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
MiroEval: Benchmarking Multimodal Deep Research Agents in Process and Outcome
Fangda Ye, Yuxin Hu, Pengxiang Zhu +19
Recent progress in deep research systems has been impressive, but evaluation still lags behind real user needs. Existing benchmarks predominantly assess final reports using fixed r…
cs.AI2026
GMP: A Benchmark for Content Moderation under Co-occurring Violations and Dynamic Rules
Houde Dong, Yifei She, Kai Ye +3
Online content moderation is essential for maintaining a healthy digital environment, and reliance on AI for this task continues to grow. Consider a user comment using national ste…