3 papers
cs.AI2026
DR-Eval: Towards Realistic and Reproducible Deep Research Evaluation
Qianqian Xie, Qingheng Xiong, He Zhu +16
Deep Research Agents (DRAs) aim to solve complex, long-horizon research tasks involving planning, retrieval, multimodal understanding, and report generation, yet their evaluation r…
cs.CV2025
ViDiC: Video Difference Captioning
Jiangtao Wu, Shihao Li, Zhaozhou Bian +7
Understanding visual differences between dynamic scenes requires the comparative perception of compositional, spatial, and temporal changes--a capability that remains underexplored…
cs.CL2024
MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language Models
Pei Wang, Yanan Wu, Zekun Wang +12
Large Language Models (LLMs) have displayed massive improvements in reasoning and decision-making skills and can hold natural conversations with users. Recently, many tool-use benc…