Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Who&When Pro: Can LLMs Really Attribute Failures in AI Agents?
Jiale Liu, Huajun Xi, Shaokun Zhang +6
Automated failure attribution uses LLMs to identify where and why agentic systems fail. As agents become more capable, their failures become subtler, making automated attribution i…
cs.AI2026
GeoAgentBench: A Dynamic Execution Benchmark for Tool-Augmented Agents in Spatial Analysis
Bo Yu, Cheng Yang, Dongyang Hou +6
The integration of Large Language Models (LLMs) into Geographic Information Systems (GIS) marks a paradigm shift toward autonomous spatial analysis. However, evaluating these LLM-b…