2 papers
cs.AI2026
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents
Dhaval C. Patel, Kaoutar El Maghraoui, Shuxin Lin +58
Agent benchmarks are growing fast, but no single benchmark touches more than four or five of the dimensions that deployment exposes. This paper aggregates the largest coordinated d…
cs.AI2026
Towards Multi-Turn Dialog Systems for Industrial Asset Operations and Maintenance
Chengrui Li, Rujing Li, Yitong Bai +1
Industrial asset operations and maintenance question answering is inherently multi-turn, iterative, and highly dependent on external tool invocation. However, the conventional plan…