Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Towards Robust Tool Use in Agents via Experience-Driven Adaptive Guidance
Can Wang, Haoran Chen, Li Yu +4
The performance bottleneck of agents is increasingly shifting from model capability to the robustness of their execution processes. Tools play a central role as the primary interfa…
cs.AI2026
From Simple QA to Deep Research: A Verifiable Benchmark Constructed through Iterative Task Evolution
Can Wang, Haoran Chen, Haowen Gao +3
Deep research benchmarks require expert-level tasks and reliable evaluation grounded in task-specific knowledge. Existing benchmarks rely heavily on expert authoring or pre-existin…
cs.AI2026
An Effective Router for Vision-Language Model Selection
Can Wang, Shengwei Wang, Bolin Zhang +2
Vision-language models (VLMs) with varying performance and resource requirements are widely deployed, making it difficult for users to select the most appropriate one among numerou…