4 papers
Towards Robust Tool Use in Agents via Experience-Driven Adaptive Guidance
Can Wang, Haoran Chen, Li Yu +4
The performance bottleneck of agents is increasingly shifting from model capability to the robustness of their execution processes. Tools play a central role as the primary interfa…
From Simple QA to Deep Research: A Verifiable Benchmark Constructed through Iterative Task Evolution
Can Wang, Haoran Chen, Haowen Gao +3
Deep research benchmarks require expert-level tasks and reliable evaluation grounded in task-specific knowledge. Existing benchmarks rely heavily on expert authoring or pre-existin…
An Effective Router for Vision-Language Model Selection
Can Wang, Shengwei Wang, Bolin Zhang +2
Vision-language models (VLMs) with varying performance and resource requirements are widely deployed, making it difficult for users to select the most appropriate one among numerou…
Plug-and-Play Performance Estimation for LLM Services without Relying on Labeled Data
Can Wang, Dianbo Sui, Hongliang Sun +3
Large Language Model (LLM) services exhibit impressive capability on unlearned tasks leveraging only a few examples by in-context learning (ICL). However, the success of ICL varies…