Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents
Dhaval C. Patel, Kaoutar El Maghraoui, Shuxin Lin +58
Agent benchmarks are growing fast, but no single benchmark touches more than four or five of the dimensions that deployment exposes. This paper aggregates the largest coordinated d…
cs.AI2026
Synthetic Scenario Generation for Evaluation of Industry 4.0 Agents
Sagar Chethan Kumar, Rohith Kanathur, Dhaval Patel +1
Industrial agent benchmarks require realistic evaluation scenarios that integrate telemetry, failure modes, maintenance records, and domain standards. However, existing benchmarks…