14 citations · 14 across the 6 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
DAREBench: Deployment-Aware and Reliable Evaluation of Models as Agents
Yu Liu, Zhilin Liu, Zhiwei Yang +7
As large language models evolve from question-answering systems into general-purpose agents, evaluation must move beyond static answer correctness to assess multimodal perception,…
cs.AI2026
Making Every Tool Call Count: Necessary Tool-Evidence Path Rewards for Agentic Vision-Language Models
Xingming Long, Yu Liu, Zhiwei Yang +7
Modern vision-language models (VLMs) can directly answer many image-grounded questions, yet they often struggle with complex queries requiring fine-grained visual details or extern…