Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers
Tianyu Huai, Tingshuo Fan, Xinchi Chen +5
As LLMs evolve from code completion systems into autonomous scientific agents, evaluating their ability to conduct experiments has become increasingly important. Existing benchmark…
cs.AI2025
VehicleWorld: A Highly Integrated Multi-Device Environment for Intelligent Vehicle Interaction
Jie Yang, Jiajun Chen, Zhangyue Yin +7
Intelligent vehicle cockpits present unique challenges for API Agents, requiring coordination across tightly-coupled subsystems that exceed typical task environments' complexity. T…
cs.AI2025
FamilyTool: A Multi-hop Personalized Tool Use Benchmark
Yuxin Wang, Yiran Guo, Yining Zheng +7
The integration of tool learning with Large Language Models (LLMs) has expanded their capabilities in handling complex tasks by leveraging external tools. However, existing benchma…